代码之家  ›  专栏  ›  技术社区  ›  Evg

python线程池问题(等等)

  •  1
  • Evg  · 技术社区  · 15 年前

    我用threadpool编写了一个简单的网站crowler。问题是:然后爬虫是得到所有的网站,它必须完成,但在现实中,它等待的东西在最后,脚本没有完成,为什么会发生这种情况?

    from Queue import Queue
    from threading import Thread
    
    import sys
    from urllib import urlopen
    from BeautifulSoup import BeautifulSoup, SoupStrainer
    import re
    from Queue import Queue, Empty
    from threading import Thread
    
    visited = set()
    queue = Queue()
    
    class Worker(Thread):
        """Thread executing tasks from a given tasks queue"""
        def __init__(self, tasks):
            Thread.__init__(self)
            self.tasks = tasks
            self.daemon = True
            self.start()
    
        def run(self):
            while True:
                func, args, kargs = self.tasks.get()
                print "startcall in thread",self
                print args
                try: func(*args, **kargs)
                except Exception, e: print e
                print "stopcall in thread",self
                self.tasks.task_done()
    
    class ThreadPool:
        """Pool of threads consuming tasks from a queue"""
        def __init__(self, num_threads):
            self.tasks = Queue(num_threads)
            for _ in range(num_threads): Worker(self.tasks)
    
        def add_task(self, func, *args, **kargs):
            """Add a task to the queue"""
            self.tasks.put((func, args, kargs))
    
        def wait_completion(self):
            """Wait for completion of all the tasks in the queue"""
            self.tasks.join()
    
    
    def process(pool,host,url):
    
        try:
            print "get url",url
            #content = urlopen(url).read().decode(charset)
            content = urlopen(url).read()
        except UnicodeDecodeError:
            return
    
        for link in BeautifulSoup(content, parseOnlyThese=SoupStrainer('a')):
            #print "link",link
            try:
                href = link['href']
            except KeyError:
                continue
    
    
            if not href.startswith('http://'):
                href = 'http://%s%s' % (host, href)
            if not href.startswith('http://%s%s' % (host, '/')):
                continue
    
    
    
            if href not in visited:
                visited.add(href)
                pool.add_task(process,pool,host,href)
                print href
    
    
    
    
    def start(host,charset):
    
        pool = ThreadPool(7)
        pool.add_task(process,pool,host,'http://%s/' % (host))
        pool.wait_completion()
    
    start('simplesite.com','utf8') 
    
    1 回复  |  直到 15 年前
        1
  •  1
  •   dugres    15 年前

    我看到的问题是你从来没有放弃过 虽然 在里面 . 所以,会的 永远。你需要在工作完成后打破这个循环。

    你可以尝试:

    if not func: break  
    

    之后 任务.get(...) 运行 .

    2) 追加

    pool.add_task(None, None, None)  
    

    结束时 .

    这是一种 过程 通知 他没有更多的任务要处理。