代码之家  ›  专栏  ›  技术社区  ›  Mike Q

InputStreamReader缓冲问题

  •  11
  • Mike Q  · 技术社区  · 16 年前

    我从一个文件中读取数据,不幸的是,这个文件有两种字符编码。

    有一个头和一个身体。标头始终为ASCII格式,并定义正文编码的字符集。

    文件也可能相当大,所以我需要避免将整个内容带到内存中。

    所以我从一个输入流开始。我首先用一个带有ASCII的InputStreamReader包装它,然后解码头并提取正文的字符集。一切都很好。

    然后我用正确的字符集创建一个新的InputStreamReader,将它放到同一个InputStream上并开始尝试读取正文。

    不幸的是,javadoc证实了这一点,InputStreamReader可能会选择提前阅读以提高效率。所以头的读数会影响身体的部分/全部。

    有人对解决这个问题有什么建议吗?手动创建一个CharsetDecoder,一次输入一个字节,这是一个好主意吗(可能包装在一个定制的阅读器实现中?)

    提前谢谢。

    // An InputStreamReader that only consumes as many bytes as is necessary
    // It does not do any read-ahead.
    public class InputStreamReaderUnbuffered extends Reader
    {
        private final CharsetDecoder charsetDecoder;
        private final InputStream inputStream;
        private final ByteBuffer byteBuffer = ByteBuffer.allocate( 1 );
    
        public InputStreamReaderUnbuffered( InputStream inputStream, Charset charset )
        {
            this.inputStream = inputStream;
            charsetDecoder = charset.newDecoder();
        }
    
        @Override
        public int read() throws IOException
        {
            boolean middleOfReading = false;
    
            while ( true )
            {
                int b = inputStream.read();
    
                if ( b == -1 )
                {
                    if ( middleOfReading )
                        throw new IOException( "Unexpected end of stream, byte truncated" );
    
                    return -1;
                }
    
                byteBuffer.clear();
                byteBuffer.put( (byte)b );
                byteBuffer.flip();
    
                CharBuffer charBuffer = charsetDecoder.decode( byteBuffer );
    
                // although this is theoretically possible this would violate the unbuffered nature
                // of this class so we throw an exception
                if ( charBuffer.length() > 1 )
                    throw new IOException( "Decoded multiple characters from one byte!" );
    
                if ( charBuffer.length() == 1 )
                    return charBuffer.get();
    
                middleOfReading = true;
            }
        }
    
        public int read( char[] cbuf, int off, int len ) throws IOException
        {
            for ( int i = 0; i < len; i++ )
            {
                int ch = read();
    
                if ( ch == -1 )
                    return i == 0 ? -1 : i;
    
                cbuf[ i ] = (char)ch;
            }
    
            return len;
        }
    
        public void close() throws IOException
        {
            inputStream.close();
        }
    }
    
    6 回复  |  直到 16 年前
        1
  •  3
  •   bruno conde    16 年前

    InputStream

    输入流 应该 skip 头字节。

        2
  •  3
  •   palacsint    14 年前

    这是伪代码。

    1. 使用 InputStream Reader 围绕着它。
    2. 将它们存储到 ByteArrayOutputStream .
    3. 创建 ByteArrayInputStream ByteArrayOutputStream公司 标题,这次换行 字节数组输入流 进入之内 读卡器 使用ASCII字符集。
    4. 输入,并读取该字节数 .
    5. 创建另一个 字节数组输入流 ByteArrayOutputStream公司 把它包起来 读卡器 带着 标题。
        3
  •  1
  •   Tom Hawtin - tackline    16 年前

    InputStreamReader . 或许可以假设 InputStream.mark 支持。

        4
  •  1
  •   T.J. Crowder    16 年前

    我的第一个想法是关闭流并重新打开它,使用 InputStream#skip InputStreamReader

    如果你真的,真的不想重新打开文件,你可以使用 file descriptors 要获取文件中的多个流,尽管您可能必须使用 channels 在文件中有多个位置(因为你不能假设你可以用 reset ,可能不支持)。

        5
  •  1
  •   derBiggi    16 年前

    更简单的是:

    private Reader reader;
    private InputStream stream;
    
    public void read() {
        int c = 0;
        while ((c = stream.read()) != -1) {
            // Read encoding
            if ( headerFullyRead ) {
                reader = new InputStreamReader( stream, encoding );
                break;
            }
        }
        while ((c = reader.read()) != -1) {
            // Handle rest of file
        }
    }
    
        6
  •  1
  •   Michael Haefele    11 年前

    如果包装InputStream并将所有读取限制为一次仅读取1字节,则似乎禁用了InputStreamReader内部的缓冲。

    这样我们就不必重写InputStreamReader逻辑。

    public class OneByteReadInputStream extends InputStream
    {
        private final InputStream inputStream;
    
        public OneByteReadInputStream(InputStream inputStream)
        {
            this.inputStream = inputStream;
        }
    
        @Override
        public int read() throws IOException
        {
            return inputStream.read();
        }
    
        @Override
        public int read(byte[] b, int off, int len) throws IOException
        {
            return super.read(b, off, 1);
        }
    }
    

    new InputStreamReader(new OneByteReadInputStream(inputStream));