代码之家  ›  专栏  ›  技术社区  ›  kzh

从python中的单词列表返回随机单词

  •  6
  • kzh  · 技术社区  · 17 年前

    我想使用python从一个文件中检索一个随机单词,但我不相信我的以下方法是最好的或有效的。请协助。

    import fileinput
    import _random
    file = [line for line in fileinput.input("/etc/dictionaries-common/words")]
    rand = _random.Random()
    print file[int(rand.random() * len(file))],
    
    9 回复  |  直到 10 年前
        1
  •  17
  •   Gourneau    12 年前

    随机模块定义choice(),它执行您想要的操作:

    import random
    
    words = [line.strip() for line in open('/etc/dictionaries-common/words')]
    print(random.choice(words))
    

    还要注意,这假设每个单词都在文件中的一行上。如果文件很大,或者经常执行此操作,您可能会发现不断重读文件会对应用程序的性能产生负面影响。

        2
  •  9
  •   Nadia Alramli    17 年前

    另一个解决方案是 getline

    import linecache
    import random
    line_number = random.randint(0, total_num_lines)
    linecache.getline('/etc/dictionaries-common/words', line_number)
    

    从文档中:

    linecache模块允许 任何文件中的任何行,而 尝试内部优化, 使用缓存,常见的情况是 从一个文件中读取许多行

    编辑: 您可以计算一次总数并存储它,因为字典文件不太可能更改。

        3
  •  9
  •   jfs    17 年前
    >>> import random
    >>> random.choice(list(open('/etc/dictionaries-common/words')))
    'jaundiced\n'
    

    它在人类时间上是有效的。

    顺便说一句,您的实现与stdlib的实现一致 random.py :

     def choice(self, seq):
        """Choose a random element from a non-empty sequence."""
        return seq[int(self.random() * len(seq))]  
    

    测量时间性能

    我想知道所提出的解决方案的相对性能是什么。 linecache -基础是最受欢迎的。慢了多少 random.choice 与在 select_random_line() ?

    # nadia_known_num_lines   9.6e-06 seconds 1.00
    # nadia                   0.056 seconds 5843.51
    # jfs                     0.062 seconds 1.10
    # dcrosta_no_strip        0.091 seconds 1.48
    # dcrosta                 0.13 seconds 1.41
    # mark_ransom_no_strip    0.66 seconds 5.10
    # mark_ransom_choose_from 0.67 seconds 1.02
    # mark_ransom             0.69 seconds 1.04
    

    (每个函数被调用10次(缓存性能))。

    这些结果表明,简单的解决方案( dcrosta )在这种情况下比故意的要快( mark_ransom )。

    用于比较的代码( as a gist ):

    import linecache
    import random
    from timeit import default_timer
    
    
    WORDS_FILENAME = "/etc/dictionaries-common/words"
    
    
    def measure(func):
        measure.func_to_measure.append(func)
        return func
    measure.func_to_measure = []
    
    
    @measure
    def dcrosta():
        words = [line.strip() for line in open(WORDS_FILENAME)]
        return random.choice(words)
    
    
    @measure
    def dcrosta_no_strip():
        words = [line for line in open(WORDS_FILENAME)]
        return random.choice(words)
    
    
    def select_random_line(filename):
        selection = None
        count = 0
        for line in file(filename, "r"):
            if random.randint(0, count) == 0:
                selection = line.strip()
                count = count + 1
        return selection
    
    
    @measure
    def mark_ransom():
        return select_random_line(WORDS_FILENAME)
    
    
    def select_random_line_no_strip(filename):
        selection = None
        count = 0
        for line in file(filename, "r"):
            if random.randint(0, count) == 0:
                selection = line
                count = count + 1
        return selection
    
    
    @measure
    def mark_ransom_no_strip():
        return select_random_line_no_strip(WORDS_FILENAME)
    
    
    def choose_from(iterable):
        """Choose a random element from a finite `iterable`.
    
        If `iterable` is a sequence then use `random.choice()` for efficiency.
    
        Return tuple (random element, total number of elements)
        """
        selection, i = None, None
        for i, item in enumerate(iterable):
            if random.randint(0, i) == 0:
                selection = item
    
        return selection, (i+1 if i is not None else 0)
    
    
    @measure
    def mark_ransom_choose_from():
        return choose_from(open(WORDS_FILENAME))
    
    
    @measure
    def nadia():
        global total_num_lines
        total_num_lines = sum(1 for _ in open(WORDS_FILENAME))
    
        line_number = random.randint(0, total_num_lines)
        return linecache.getline(WORDS_FILENAME, line_number)
    
    
    @measure
    def nadia_known_num_lines():
        line_number = random.randint(0, total_num_lines)
        return linecache.getline(WORDS_FILENAME, line_number)
    
    
    @measure
    def jfs():
        return random.choice(list(open(WORDS_FILENAME)))
    
    
    def timef(func, number=1000, timer=default_timer):
        """Return number of seconds it takes to execute `func()`."""
        start = timer()
        for _ in range(number):
            func()
        return (timer() - start) / number
    
    
    def main():
        # measure time
        times = dict((f.__name__, timef(f, number=10))
                     for f in measure.func_to_measure)
    
        # print from fastest to slowest
        maxname_len = max(map(len, times))
        last = None
        for name in sorted(times, key=times.__getitem__):
            print "%s %4.2g seconds %.2f" % (name.ljust(maxname_len), times[name],
                                             last and times[name] / last or 1)
            last = times[name]
    
    
    if __name__ == "__main__":
        main()
    
        4
  •  3
  •   Community Mohan Dere    9 年前

    我的答案来自 What’s the best way to return a random line in a text file using C? 以下内容:

    import random
    
    def select_random_line(filename):
        selection = None
        count = 0
        for line in file(filename, "r"):
            if random.randint(0, count) == 0:
                selection = line.strip()
            count = count + 1
        return selection
    
    print select_random_line("/etc/dictionaries-common/words")
    

    编辑:使用的原始答案版本 readlines 这并不像我想的那样有效,完全没有必要。这个版本将迭代文件,而不是将其全部读取到内存中,并在一次传递中完成,这将使它比我迄今为止看到的任何答案都更有效率。

    广义版本

    import random
    
    def choose_from(iterable):
        """Choose a random element from a finite `iterable`.
    
        If `iterable` is a sequence then use `random.choice()` for efficiency.
    
        Return tuple (random element, total number of elements)
        """
        selection, i = None, None
        for i, item in enumerate(iterable):
            if random.randint(0, i) == 0:
                selection = item
    
        return selection, (i+1 if i is not None else 0)
    

    实例

    print choose_from(open("/etc/dictionaries-common/words"))
    print choose_from(dict(a=1, b=2))
    print choose_from(i for i in range(10) if i % 3 == 0)
    print choose_from(i for i in range(10) if i % 11 == 0 and i) # empty
    print choose_from([0]) # one element
    chunk, n = choose_from(urllib2.urlopen("http://google.com"))
    print (chunk[:20], n)
    

    产量

    ('yeps\n', 98569)
    ('a', 2)
    (6, 4)
    (None, 0)
    (0, 1)
    ('window._gjp && _gjp(', 10)
    
        5
  •  2
  •   Dwight Kelly    17 年前
        6
  •  1
  •   Greg Hewgill    17 年前

    你可以不用 fileinput :

    import random
    data = open("/etc/dictionaries-common/words").readlines()
    print random.choice(data)
    

    我也用过 data 而不是 file 因为 文件 是Python中的预定义类型。

        7
  •  1
  •   Jason Christa    17 年前

    我没有你的代码,但就算法而言:

    1. 查找文件大小
    2. 使用seek()函数执行随机搜索
    3. 查找下一个(或上一个)空白字符
    4. 返回在该空白字符之后开始的单词
        8
  •  0
  •   Oli    17 年前

    在这种情况下,效率和冗长是不一样的。这是一种非常诱人的方法,它可以用一行或两行的方式完成所有工作,但是对于文件I/O,要坚持经典的fopen风格,低级别的交互,即使它需要更多的代码行。

    我可以复制和粘贴一些代码,并声称它是我自己的(其他人可以,如果他们愿意的话),但看看这个: http://mail.python.org/pipermail/tutor/2007-July/055635.html

        9
  •  0
  •   pafcu    17 年前

    有几种不同的方法来优化这个问题。您可以优化速度或空间。

    如果您想要一个快速但需要内存的解决方案,请使用file.readlines()读取整个文件,然后使用random.choice()。

    如果您想要一个内存高效的解决方案,首先通过反复调用somefile.readline()来检查文件中的行数,直到它返回“”,然后生成一个小于行数(例如,n)的随机数,返回到文件的开头,最后调用somefile.readline()n次。下次调用somefile.readline()将返回所需的随机行。这种方法不浪费内存来保存“不必要的”行。当然,如果您计划从文件中获取大量随机行,这将是非常低效的,而且最好将整个文件保存在内存中,就像第一种方法一样。