代码之家  ›  专栏  ›  技术社区  ›  tourist

如何在pyspark中读取二进制数据

  •  0
  • tourist  · 技术社区  · 6 年前

    我正在读二进制文件 http://snap.stanford.edu/data/amazon/productGraph/image_features/image_features.b

    import array
    from io import StringIO
    
    img_embedding_file = sc.binaryRecords("s3://bucket/image_features.b", 4106)
    
    def mapper(features):
        a = array.array('f')
        a.frombytes(features)
        return a.tolist()
    
    def byte_mapper(bytes):
        return str(bytes)
    
    decoded_embeddings = img_embedding_file.map(lambda x: [byte_mapper(x[:10]), mapper(x[10:])])
    
    

    什么时候 product_id 使用从rdd中选择

    decoded_embeddings = img_embedding_file.map(lambda x: [byte_mapper(x[:10]), mapper(x[10:])])
    

    输出 是

    ["b'1582480311'", "b'\\x00\\x00\\x00\\x00\\x88c-?\\xeb\\xe2'", "b'7@\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00'", "b'\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00'", "b'\\xec/\\x0b?\\x00\\x00\\x00\\x00K\\xea'", "b'\\x00\\x00c\\x7f\\xd9?\\x00\\x00\\x00\\x00'", "b'L\\xa6\\n>\\x00\\x00\\x00\\x00\\xfe\\xd4'", "b'\\x00\\x00\\x00\\x00\\x00\\x00\\xe5\\xd0\\xa2='", "b'\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00'", "b'\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00\\x00'"]
    

    该文件位于s3上。 每行中的文件都有前10个字节 image_features 我能够提取所有4096图像特征,但在读取前10个字节并将其转换为适当的可读格式时面临问题。

    0 回复  |  直到 6 年前
        1
  •  1
  •   blackbishop    6 年前

    编辑

    recordLength . 不是 4096 + 10 但是 4096*4 + 10 . 转向:

    img_embedding_file = sc.binaryRecords("s3://bucket/image_features.b", 16394)
    

    应该有用。 provided code 您从网站下载了二进制文件:

    for i in range(4096):
         feature.append(struct.unpack('f', f.read(4))) # <-- so 4096 * 4
    

    :

    我认为这个问题来自你的建议 byte_mapper 这不是将字节转换为字符串的正确方法。你应该使用 decode

    bytes = b'1582480311'
    print(str(bytes))
    # output: "b'1582480311'"
    
    print(bytes.decode("utf-8"))
    # output: '1582480311'
    

    如果出现错误:

    UnicodeDecodeError:“utf-8”编解码器无法解码位置4中的字节0x88:无效的开始字节

    这意味着 product_id

    但是,您可能希望通过添加选项来忽略这些字符 ignore 到 解码 功能:

    bytes.decode("utf-8", "ignore")