代码之家  ›  专栏  ›  技术社区  ›  Programmer_nltk

百分比计数动词,名词使用空格?

  •  0
  • Programmer_nltk  · 技术社区  · 8 年前

    我想用spacy计算一个句子中POS的百分比,类似于

    Count verbs, nouns, and other parts of speech with python's NLTK

    目前能够检测和计算位置。如何找到百分比分割。

    from __future__ import unicode_literals
    import spacy,en_core_web_sm
    from collections import Counter
    nlp = en_core_web_sm.load()
    print Counter(([token.pos_ for token in nlp('The cat sat on the mat.')]))
    

    电流输出:

    Counter({u'NOUN': 2, u'DET': 2, u'VERB': 1, u'ADP': 1, u'PUNCT': 1})
    

    预期产量:

    Noun: 28.5%
    DET: 28.5%
    VERB: 14.28%
    ADP: 14.28%
    PUNCT: 14.28%
    

    如何将输出写入pandas数据帧?

    2 回复  |  直到 7 年前
        1
  •  1
  •   sophros    8 年前

    这些方面的东西应该能满足你的需要:

    sbase = sum(c.values())
    
    for el, cnt in c.items():
        print(el, '{0:2.2f}%'.format((100.0* cnt)/sbase))
    
    
    NOUN 28.57%
    DET 28.57%
    VERB 14.29%
    ADP 14.29%
    PUNCT 14.29%
    
        2
  •  0
  •   Programmer_nltk    8 年前
    from __future__ import unicode_literals
    import spacy,en_core_web_sm
    from collections import Counter
    nlp = en_core_web_sm.load()
    c = Counter(([token.pos_ for token in nlp('The cat sat on the mat.')]))
    sbase = sum(c.values())
    for el, cnt in c.items():
        print(el, '{0:2.2f}%'.format((100.0* cnt)/sbase))
    

    输出:

    (u'NOUN', u'28.57%')
    (u'VERB', u'14.29%')
    (u'DET', u'28.57%')
    (u'ADP', u'14.29%')
    (u'PUNCT', u'14.29%')