代码之家  ›  专栏  ›  技术社区  ›  Paul R

谷歌语音API流音频

  •  2
  • Paul R  · 技术社区  · 8 年前

    在文档页中 https://cloud.google.com/speech/ 有一个演示示例通过浏览器收听语音并在后台使用api。这个演示的源代码可用吗?

    2 回复  |  直到 8 年前
        1
  •  1
  •   Sunil    8 年前

    google cloud speech页面上的演示并没有直接从浏览器中使用api。他们打开一个websocket到一个后端服务器,该服务器实际上与来自服务器的语音api进行对话。 可以在客户端浏览器上直接使用 REST api 但是,如果您想要实时转录,您必须拥有自己的中介服务器,其中包含websocket中介到 gRPC based Speech api

    您可以在此处找到基于websocket的体系结构的代码/演示: https://github.com/googlecodelabs/speaking-with-a-webpage

        2
  •  0
  •   Dawn T Cherian    7 年前

    对于python websocket服务器,您可以遵循以下步骤 code

    import asyncio
    import websockets
    import json
    import threading
    from six.moves import queue
    from google.cloud import speech
    from google.cloud.speech import types
    
    
    IP = '0.0.0.0'
    PORT = 8000
    
    class Transcoder(object):
        """
        Converts audio chunks to text
        """
        def __init__(self, encoding, rate, language):
            self.buff = queue.Queue()
            self.encoding = encoding
            self.language = language
            self.rate = rate
            self.closed = True
            self.transcript = None
    
        def start(self):
            """Start up streaming speech call"""
            threading.Thread(target=self.process).start()
    
        def response_loop(self, responses):
            """
            Pick up the final result of Speech to text conversion
            """
            for response in responses:
                if not response.results:
                    continue
                result = response.results[0]
                if not result.alternatives:
                    continue
                transcript = result.alternatives[0].transcript
                if result.is_final:
                    self.transcript = transcript
    
        def process(self):
            """
            Audio stream recognition and result parsing
            """
            #You can add speech contexts for better recognition
            cap_speech_context = types.SpeechContext(phrases=["Add your phrases here"])
            client = speech.SpeechClient()
            config = types.RecognitionConfig(
                encoding=self.encoding,
                sample_rate_hertz=self.rate,
                language_code=self.language,
                speech_contexts=[cap_speech_context,],
                model='command_and_search'
            )
            streaming_config = types.StreamingRecognitionConfig(
                config=config,
                interim_results=False,
                single_utterance=False)
            audio_generator = self.stream_generator()
            requests = (types.StreamingRecognizeRequest(audio_content=content)
                        for content in audio_generator)
    
            responses = client.streaming_recognize(streaming_config, requests)
            try:
                self.response_loop(responses)
            except:
                self.start()
    
        def stream_generator(self):
            while not self.closed:
                chunk = self.buff.get()
                if chunk is None:
                    return
                data = [chunk]
                while True:
                    try:
                        chunk = self.buff.get(block=False)
                        if chunk is None:
                            return
                        data.append(chunk)
                    except queue.Empty:
                        break
                yield b''.join(data)
    
        def write(self, data):
            """
            Writes data to the buffer
            """
            self.buff.put(data)
    
    
    async def audio_processor(websocket, path):
        """
        Collects audio from the stream, writes it to buffer and return the output of Google speech to text
        """
        config = await websocket.recv()
        if not isinstance(config, str):
            print("ERROR, no config")
            return
        config = json.loads(config)
        transcoder = Transcoder(
            encoding=config["format"],
            rate=config["rate"],
            language=config["language"]
        )
        transcoder.start()
        while True:
            try:
                data = await websocket.recv()
            except websockets.ConnectionClosed:
                print("Connection closed")
                break
            transcoder.write(data)
            transcoder.closed = False
            if transcoder.transcript:
                print(transcoder.transcript)
                await websocket.send(transcoder.transcript)
                transcoder.transcript = None
    
    start_server = websockets.serve(audio_processor, IP, PORT)
    asyncio.get_event_loop().run_until_complete(start_server)
    asyncio.get_event_loop().run_forever()