代码之家  ›  专栏  ›  技术社区  ›  TofuBeer

从长度为无符号整数的ByteBuffer读取UTF-8字符串

  •  1
  • TofuBeer  · 技术社区  · 17 年前

    我试图通过java.nio.ByteBuffer读取UTF8字符串。大小是一个无符号的int,当然,Java没有。我已将该值读入一个long中,以便获得该值。

    我遇到的下一个问题是,我无法创建一个长字节数组,将长字节转换回int将导致它被签名。

    我还尝试在缓冲区上使用limit(),但它同样适用于int-not-long。

    关于如何从ByteBuffer读取可能长度为无符号int的UTF8字符串的任何想法。

    编辑:

    Here is an example of the issue .

    SourceDebugExtension_attribute {
           u2 attribute_name_index;
           u4 attribute_length;
           u1 debug_extension[attribute_length];
        }
    
    attribute_name_index
        The value of the attribute_name_index item must be a valid index into the constant_pool table. The constant_pool entry at that index must be a CONSTANT_Utf8_info structure representing the string "SourceDebugExtension".
    
    attribute_length
        The value of the attribute_length item indicates the length of the attribute, excluding the initial six bytes. The value of the attribute_length item is thus the number of bytes in the debug_extension[] item.
    
    debug_extension[]
        The debug_extension array holds a string, which must be in UTF-8 format. There is no terminating zero byte.
    
        The string in the debug_extension item will be interpreted as extended debugging information. The content of this string has no semantic effect on the Java Virtual Machine.
    

    因此,从技术角度来看,类文件中可能有一个长度为完整u4(无符号,4字节)的字符串。

    如果UTF8字符串的大小有限制(我不是UTF8专家,所以可能有这样的限制),那么这些不会成为问题。

    4 回复  |  直到 14 年前
        1
  •  6
  •   Alnitak    17 年前

    除非您的字节数组大于2GB(Java的最大正值) int ),您将不会有铸造的问题 long 回到一个签名的 int

    如果您的字节数组长度需要超过2GB,那么您就做错了,尤其是因为这远远超过了JVM的默认最大堆化。。。

        2
  •  1
  •   Peter Lawrey    17 年前

    签名int不会是你的主要问题。假设你有一根40亿长的绳子。您需要一个至少为4GB的字节缓冲区,一个至少为4GB的字节[]。将其转换为字符串时,至少需要8GB(每个字符2字节)和StringBuilder来构建它。(至少为8 GB) 所有您需要的,24 GB的处理1字符串。即使你有很多内存,你也不会得到很多这样大小的字符串。

    另一种方法是将长度视为有符号,如果无符号,则视为错误,因为在任何情况下都没有足够的内存来处理字符串。即使要处理一个长度为20亿(2^31-1)的字符串,您也需要12 GB才能以这种方式将其转换为字符串。

        3
  •  1
  •   Michael Borgwardt    17 年前

    as per the languge spec

    但是,即使在一个块中处理这么多也太多了——如果遇到足够大的字符串,它将完全破坏性能,并使您的程序在大多数机器上失败,并出现OutOfMemory错误。

    您应该做的是将任何字符串处理成合理大小的块,比如说一次处理几兆字节。那么,您可以处理的大小没有实际限制。

        4
  •  0
  •   Wilfred Springer    16 年前

    我想你可以实施 CharSequence 在一辆小货车上。这将允许您防止“字符串”出现在堆上,尽管大多数处理字符的实用程序实际上都需要字符串。即使这样,字符序列也有一个限制。它希望大小以int形式返回。

    subSequence(...) 返回普通字符序列。)