{
    "componentChunkName": "component---src-templates-post-tsx",
    "path": "/audio_app/",
    "result": {"data":{"logo":null,"markdownRemark":{"html":"<h1>AI &#x26; Speech Processing: Application-1</h1>\n<p>본 글은 광운대학교 전자공학과 박호종 교수님의 강의를 듣고 작성되었음을 밝힙니다.</p>\n<h2>음성/오디오/sound 특성</h2>\n<ul>\n<li>방송 / 통신 / 엔터테인먼트의 핵심 기술</li>\n<li>학계/산업계의 전문 엔지니어 부족</li>\n<li>Art, 취미 활동과 관련된 기술</li>\n<li>음성에서 언어의 종속성</li>\n</ul>\n<p>음성 이해, 음성 신호의 품질 및 명료도 평가에서 중요한 요인</p>\n<ul>\n<li>음성은 대표적인 생체신호</li>\n</ul>\n<p>음성 기반 헬스케어, 장애인을 위한 복지 기술</p>\n<ul>\n<li>Data 양 : V >> A (4차원 : 2차원)</li>\n<li>심리적 민감도 : V &#x3C; A</li>\n</ul>\n<p>특정 sound에 대한 거부감 존재<br>\n심리 변화에 큰 영향 : 공포 영화에서 분위기 조정</p>\n<ul>\n<li>요구 품질 : V &#x3C; A</li>\n</ul>\n<p>일반적으로 낮은 화질은 허용되지만 낮은 음질은 허용되지 않음</p>\n<ul>\n<li>멀티미디어에서의 독립성 : V &#x3C; A</li>\n</ul>\n<h2>Applications - Speech Synthesis</h2>\n<p>기계를 이용하여 언어 정보를 가지는 음성 신호를 생성 : Text-To-Speech (TTS)</p>\n<p><img src=\"https://user-images.githubusercontent.com/42150335/78545239-6f5aa500-7836-11ea-8851-b87875dd586f.png\" alt=\"tts\"></p>\n<p>음정을 맞추기는 쉽지만, 발음을 정확히 맞추기가 어렵다.</p>\n<ul>\n<li>Before 2010 : digital waveform 연결, boundary smoothing</li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/42150335/78545482-cfe9e200-7836-11ea-93e1-dc07293a672c.png\" alt=\"tts-before-2010\"></p>\n<p>음절 단위로 미리 녹음해놓고 파형을 저장해놓는다.<br>\nText를 보고 미리 녹음해놓은 파형을 적절하게 이어붙인다.</p>\n<p>=> 부자연스러움</p>\n<p>이러한 부자연스러움을 해결하기 위해 단어 단위로 녹음하는 것이 더 자연스러웠음</p>\n<p>하지만 자연스럽게 하는 과정이 굉장히 어렵다.</p>\n<p>Example)</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">고기\n\n소고기\n돼지고기\n\n불고기\n\n물고기\n=> 물꼬기라고 발음이 됨 ※ 문제가 됨 ※</code></pre></div>\n<ul>\n<li>After 2010 : AI-based waveform generation</li>\n</ul>\n<p>현재 AI를 이용하여 상용화가 가능할 정도의 음질이 나오게 됨</p>\n<ul>\n<li>부모가 들려주는 동화책</li>\n<li>죽은 사람의 목소리 재현 : 신체 구조로부터 음색 추정</li>\n</ul>\n<h2>Applications - Music Synthesis</h2>\n<p>전자장치를 이용하여 music signal 합성\n(쉽게 말하면 전자 키보드)</p>\n<h2>Applications - Sound for Game and Animation</h2>\n<p>게임이나 애니메이션 사운드를 직접 만드는 기술<br>\n(물건이 떨어지는 소리, 발자국 소리 등..)</p>\n<ul>\n<li>기존에는 미리 녹음된 waveform을 상황에 맞추어 출력</li>\n</ul>\n<p>=> 단순한 sound 반복에 의한 피로감</p>\n<ul>\n<li>Sound 합성 엔진 이용</li>\n</ul>\n<p>물체 특성과 움직임에 따라 수학적으로 sound를 합성</p>\n<h2>Applications - Speech Recognition</h2>\n<p>기계가 음성 신호에 포함되어 있는 언어정보를 인식</p>\n<p><img src=\"https://user-images.githubusercontent.com/42150335/78549178-ebf08200-783c-11ea-9dfa-3922f0d77404.png\" alt=\"stt\"></p>\n<ul>\n<li>Human-machine interface의 핵심 기술</li>\n</ul>\n<p>휴대전화에서 mic 입력으로 문자 및 명령 입력<br>\n장애인, 특수 환경에서의 기기 동작<br>\n지능형 로봇</p>\n<ul>\n<li>언어에 대한 지식 필요</li>\n<li>감정 인식에 대한 연구 진행</li>\n<li>문제점</li>\n</ul>\n<p>기술적 한계 : 사용자는 매우 높은 인식률 요구<br>\n사용에 대한 거부감 (의외로 불편)<br>\n더 편리한 다른 방법이 있으면 사용하지 않음<br>\n=> 말보다 키보드 or 마우스 입력이 편하다</p>\n<ul>\n<li>자연어 (Natural Language) 인식</li>\n</ul>\n<p>휴대폰에서 많이 사용하는 이유는, 조그마한 휴대폰에 타이핑하기가 힘들기 때문이다.<br>\n음성인식은 항상 첫번째 옵션이 아닌, 두번째 세번째 옵션인 것을 이해해야 한다.</p>\n<h2>AI Assistant</h2>\n<p>아마존의 에코, SK의 누구, KT의 기가지니 등 최근 많이 볼 수 있는 제품</p>\n<p>생각보다 불편한 탓에 아직 활용가치가 높지 않다.</p>\n<ul>\n<li>문제점 1</li>\n</ul>\n<p>아마존 인형의집 주문 사건</p>\n<p>TV에서 나온 소리를 인식해서, 미국 전역에 장난감을 주문한 사건</p>\n<ul>\n<li>문제점 2</li>\n</ul>\n<p>화자 인식</p>\n<p><img src=\"https://user-images.githubusercontent.com/42150335/78549804-e8112f80-783d-11ea-83fd-3f6bb63870c4.png\" alt=\"assistant-problem\"></p>\n<p>=> 여기서 “내”가 누군지 모른다.</p>\n<ul>\n<li>Privacy Issue</li>\n</ul>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">항상 Sound를 수집하고 있다?  \n  \n수집은 하지만 폐기한다?  \n  \n일상 대화  \n  \n인간 인식</code></pre></div>\n<p>Camera (CCTV)에 비하여 위험성은 인지 못함</p>\n<h2>자동차에서의 음성 인식</h2>\n<p>자동차 내에서는 “복잡한 입력보다는 음성으로 입력을 하자”라는 논리가 성립이 됨</p>\n<p>But! 자동차라는 이유로 생기는 문제점이 있음</p>\n<ul>\n<li>다양한 잡음 (자동차 잡음, 라디오 소리, 대화 소리 등 …)</li>\n<li>원거리 마이크</li>\n<li>버튼보다 불편함</li>\n</ul>\n<h2>Very Efficient Speech Communication</h2>\n<ul>\n<li>파형 전송 없이 텍스트만 전송하므로 정보량이 매우 적음</li>\n<li>자연스러운 통신을 위하여 감정과 Speaker 특성 전송</li>\n</ul>\n<p><img src=\"https://user-images.githubusercontent.com/42150335/78550813-b8fbbd80-783f-11ea-9ad4-c97440914b0c.png\" alt=\"image\"></p>\n<p>여기에 기계 번역까지 더해진다면, 어느 언어와도 편리한 통신이 가능함</p>\n<p>아직은 Ideal한 얘기지만, 이렇게 될 것이다.</p>","htmlAst":{"type":"root","children":[{"type":"element","tagName":"h1","properties":{},"children":[{"type":"text","value":"AI & Speech Processing: Application-1"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"본 글은 광운대학교 전자공학과 박호종 교수님의 강의를 듣고 작성되었음을 밝힙니다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"음성/오디오/sound 특성"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"방송 / 통신 / 엔터테인먼트의 핵심 기술"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"학계/산업계의 전문 엔지니어 부족"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"Art, 취미 활동과 관련된 기술"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"음성에서 언어의 종속성"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"음성 이해, 음성 신호의 품질 및 명료도 평가에서 중요한 요인"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"음성은 대표적인 생체신호"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"음성 기반 헬스케어, 장애인을 위한 복지 기술"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"Data 양 : V >> A (4차원 : 2차원)"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"심리적 민감도 : V < A"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"특정 sound에 대한 거부감 존재"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n심리 변화에 큰 영향 : 공포 영화에서 분위기 조정"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"요구 품질 : V < A"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"일반적으로 낮은 화질은 허용되지만 낮은 음질은 허용되지 않음"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"멀티미디어에서의 독립성 : V < A"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"Applications - Speech Synthesis"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"기계를 이용하여 언어 정보를 가지는 음성 신호를 생성 : Text-To-Speech (TTS)"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"element","tagName":"img","properties":{"src":"https://user-images.githubusercontent.com/42150335/78545239-6f5aa500-7836-11ea-8851-b87875dd586f.png","alt":"tts"},"children":[]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"음정을 맞추기는 쉽지만, 발음을 정확히 맞추기가 어렵다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"Before 2010 : digital waveform 연결, boundary smoothing"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"element","tagName":"img","properties":{"src":"https://user-images.githubusercontent.com/42150335/78545482-cfe9e200-7836-11ea-93e1-dc07293a672c.png","alt":"tts-before-2010"},"children":[]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"음절 단위로 미리 녹음해놓고 파형을 저장해놓는다."},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\nText를 보고 미리 녹음해놓은 파형을 적절하게 이어붙인다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"=> 부자연스러움"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"이러한 부자연스러움을 해결하기 위해 단어 단위로 녹음하는 것이 더 자연스러웠음"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"하지만 자연스럽게 하는 과정이 굉장히 어렵다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"Example)"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"div","properties":{"className":["gatsby-highlight"],"dataLanguage":"text"},"children":[{"type":"element","tagName":"pre","properties":{"className":["language-text"]},"children":[{"type":"element","tagName":"code","properties":{"className":["language-text"]},"children":[{"type":"text","value":"고기\n\n소고기\n돼지고기\n\n불고기\n\n물고기\n=> 물꼬기라고 발음이 됨 ※ 문제가 됨 ※"}]}]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"After 2010 : AI-based waveform generation"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"현재 AI를 이용하여 상용화가 가능할 정도의 음질이 나오게 됨"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"부모가 들려주는 동화책"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"죽은 사람의 목소리 재현 : 신체 구조로부터 음색 추정"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"Applications - Music Synthesis"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"전자장치를 이용하여 music signal 합성\n(쉽게 말하면 전자 키보드)"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"Applications - Sound for Game and Animation"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"게임이나 애니메이션 사운드를 직접 만드는 기술"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n(물건이 떨어지는 소리, 발자국 소리 등..)"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"기존에는 미리 녹음된 waveform을 상황에 맞추어 출력"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"=> 단순한 sound 반복에 의한 피로감"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"Sound 합성 엔진 이용"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"물체 특성과 움직임에 따라 수학적으로 sound를 합성"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"Applications - Speech Recognition"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"기계가 음성 신호에 포함되어 있는 언어정보를 인식"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"element","tagName":"img","properties":{"src":"https://user-images.githubusercontent.com/42150335/78549178-ebf08200-783c-11ea-9dfa-3922f0d77404.png","alt":"stt"},"children":[]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"Human-machine interface의 핵심 기술"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"휴대전화에서 mic 입력으로 문자 및 명령 입력"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n장애인, 특수 환경에서의 기기 동작"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n지능형 로봇"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"언어에 대한 지식 필요"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"감정 인식에 대한 연구 진행"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"문제점"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"기술적 한계 : 사용자는 매우 높은 인식률 요구"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n사용에 대한 거부감 (의외로 불편)"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n더 편리한 다른 방법이 있으면 사용하지 않음"},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n=> 말보다 키보드 or 마우스 입력이 편하다"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"자연어 (Natural Language) 인식"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"휴대폰에서 많이 사용하는 이유는, 조그마한 휴대폰에 타이핑하기가 힘들기 때문이다."},{"type":"element","tagName":"br","properties":{},"children":[]},{"type":"text","value":"\n음성인식은 항상 첫번째 옵션이 아닌, 두번째 세번째 옵션인 것을 이해해야 한다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"AI Assistant"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"아마존의 에코, SK의 누구, KT의 기가지니 등 최근 많이 볼 수 있는 제품"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"생각보다 불편한 탓에 아직 활용가치가 높지 않다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"문제점 1"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"아마존 인형의집 주문 사건"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"TV에서 나온 소리를 인식해서, 미국 전역에 장난감을 주문한 사건"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"문제점 2"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"화자 인식"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"element","tagName":"img","properties":{"src":"https://user-images.githubusercontent.com/42150335/78549804-e8112f80-783d-11ea-83fd-3f6bb63870c4.png","alt":"assistant-problem"},"children":[]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"=> 여기서 “내”가 누군지 모른다."}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"Privacy Issue"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"div","properties":{"className":["gatsby-highlight"],"dataLanguage":"text"},"children":[{"type":"element","tagName":"pre","properties":{"className":["language-text"]},"children":[{"type":"element","tagName":"code","properties":{"className":["language-text"]},"children":[{"type":"text","value":"항상 Sound를 수집하고 있다?  \n  \n수집은 하지만 폐기한다?  \n  \n일상 대화  \n  \n인간 인식"}]}]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"Camera (CCTV)에 비하여 위험성은 인지 못함"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"자동차에서의 음성 인식"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"자동차 내에서는 “복잡한 입력보다는 음성으로 입력을 하자”라는 논리가 성립이 됨"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"But! 자동차라는 이유로 생기는 문제점이 있음"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"다양한 잡음 (자동차 잡음, 라디오 소리, 대화 소리 등 …)"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"원거리 마이크"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"버튼보다 불편함"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"h2","properties":{},"children":[{"type":"text","value":"Very Efficient Speech Communication"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"ul","properties":{},"children":[{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"파형 전송 없이 텍스트만 전송하므로 정보량이 매우 적음"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"li","properties":{},"children":[{"type":"text","value":"자연스러운 통신을 위하여 감정과 Speaker 특성 전송"}]},{"type":"text","value":"\n"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"element","tagName":"img","properties":{"src":"https://user-images.githubusercontent.com/42150335/78550813-b8fbbd80-783f-11ea-9ad4-c97440914b0c.png","alt":"image"},"children":[]}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"여기에 기계 번역까지 더해진다면, 어느 언어와도 편리한 통신이 가능함"}]},{"type":"text","value":"\n"},{"type":"element","tagName":"p","properties":{},"children":[{"type":"text","value":"아직은 Ideal한 얘기지만, 이렇게 될 것이다."}]}],"data":{"quirksMode":false}},"excerpt":"AI & Speech Processing: Application-1 본 글은 광운대학교 전자공학과 박호종 교수님의 강의를 듣고 작성되었음을 밝힙니다. 음성/오디오/sound…","fields":{"readingTime":{"text":"7 min read"}},"frontmatter":{"title":"Sooftware Speech - AI & Speech Processing: Application-1","userDate":"15 April 2020","date":"2020-04-15T10:00:00.000Z","tags":["speech"],"excerpt":null,"image":{"childImageSharp":{"gatsbyImageData":{"layout":"fullWidth","backgroundColor":"#f8f8f8","images":{"fallback":{"src":"/static/ea340e308aee7cbbeae7dcadb12a5f72/fee67/audio_app1.png","srcSet":"/static/ea340e308aee7cbbeae7dcadb12a5f72/ae605/audio_app1.png 750w,\n/static/ea340e308aee7cbbeae7dcadb12a5f72/fee67/audio_app1.png 841w","sizes":"100vw"},"sources":[{"srcSet":"/static/ea340e308aee7cbbeae7dcadb12a5f72/1211a/audio_app1.webp 750w,\n/static/ea340e308aee7cbbeae7dcadb12a5f72/d21da/audio_app1.webp 841w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.35434007134363854}}},"author":[{"id":"Soohwan Kim","bio":"Co-founder/A.I. engineer at TUNiB.","avatar":{"children":[{"gatsbyImageData":{"layout":"fullWidth","backgroundColor":"#282838","images":{"fallback":{"src":"/static/a9e6b445142b247ee4cfa66155398bb2/0d6f4/soohwan.png","srcSet":"/static/a9e6b445142b247ee4cfa66155398bb2/248f9/soohwan.png 40w,\n/static/a9e6b445142b247ee4cfa66155398bb2/fd435/soohwan.png 80w,\n/static/a9e6b445142b247ee4cfa66155398bb2/0d6f4/soohwan.png 120w","sizes":"100vw"},"sources":[{"srcSet":"/static/a9e6b445142b247ee4cfa66155398bb2/e7f45/soohwan.webp 40w,\n/static/a9e6b445142b247ee4cfa66155398bb2/589ec/soohwan.webp 80w,\n/static/a9e6b445142b247ee4cfa66155398bb2/71a38/soohwan.webp 120w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.6833333333333333}}]}}]}},"relatedPosts":{"totalCount":20,"edges":[{"node":{"id":"fa9e8cbb-841a-516f-9df6-4be0336b56b0","excerpt":"한국어 Tacotron2 이번 포스팅에서는 Tacotron2 아키텍처로 한국어 TTS 시스템을 만드는 방법에 대해 다루겠습니다. Tacotron2 Tacotron2는 17년 12월 구글이 NATURAL TTS SYNTHESIS BY…","frontmatter":{"title":"Sooftware Speech - 한국어 Tacotron2","date":"2021-10-10T10:00:00.000Z"},"fields":{"readingTime":{"text":"11 min read"},"slug":"/korean_tacotron2/"}}},{"node":{"id":"43c23529-71b1-5d60-8883-a45fbcd55ebd","excerpt":"Textless NLP: Generating expressive speech from raw audio paper / code / pre-train model / blog Name: Generative Spoken Language Model (GSLM…","frontmatter":{"title":"Sooftware NLP - Textless NLP","date":"2021-09-19T10:00:00.000Z"},"fields":{"readingTime":{"text":"4 min read"},"slug":"/Textledd NLP: Generating expressive speech from raw audio/"}}},{"node":{"id":"83c6b4fa-d71e-51d8-90bb-b58bfffc01d0","excerpt":"Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition Yu Zhang et al., 2020 Google Research, Brain Team Reference…","frontmatter":{"title":"Sooftware Speech - Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition Paper Review","date":"2021-03-17T10:00:00.000Z"},"fields":{"readingTime":{"text":"3 min read"},"slug":"/Pushing the Limits of Semi-Supervised Learning for Automatic Speech Recognition/"}}},{"node":{"id":"894a24ca-fff9-5884-b6da-1c97f2ece7bc","excerpt":"PORORO Text-To-Speech (TTS) 얼마전에 저희 팀에서 공개한 PORORO: Platform Of neuRal mOdels for natuRal language prOcessing 라이브러리에 제가 공들여만든 TTS…","frontmatter":{"title":"PORORO Text-To-Speech (TTS)","date":"2021-02-16T10:00:00.000Z"},"fields":{"readingTime":{"text":"1 min read"},"slug":"/pororo-tts/"}}},{"node":{"id":"b039977c-cecb-50f4-a0dd-fd008731bc99","excerpt":"EMNLP Paper Review: Speech Adaptive Feature Selection for End-to-End Speech Translation (Biao Zhang et al) Incremental Text-to-Speech…","frontmatter":{"title":"Sooftware Speech - EMNLP Paper Review: Speech","date":"2020-12-08T10:00:00.000Z"},"fields":{"readingTime":{"text":"4 min read"},"slug":"/2020 EMNLP Speech Paper Review/"}}}]}},"pageContext":{"slug":"/audio_app/","prev":{"excerpt":"AI & Speech Signal Processing Lecture : DSP for Audio 본 글은 광운대학교 전자공학과 박호종 교수님의 강의를 듣고 작성되었음을 밝힙니다. 이제는 오디오에 특화된 DSP로 넘어가보자. Short-Time…","frontmatter":{"title":"Sooftware Speech - AI & Speech Processing: DSP for Audio","tags":["speech","dsp"],"date":"2020-04-11T10:00:00.000Z","draft":false,"excerpt":null,"image":{"childImageSharp":{"gatsbyImageData":{"layout":"fullWidth","placeholder":{"fallback":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAATCAYAAACQjC21AAAACXBIWXMAAAsTAAALEwEAmpwYAAACrUlEQVQ4y32UvW4TQRSF/TI0dCAhJAo6aioeBImSN6DgASiQkBAooqFIAekIiVAiQgI4CcqP4xDbu7bXu97f2Z+5H5qx17GN4Uqj9czOPT7nnnu3ASAiXIcQ5sL+UPg5Er4MND9GwrEvXITCwWhyvtkVjsbCaSDoabrBaZgfeZ7zdW+P5tGxfbF5JdxfE269gjtvhNuvhQfvNA/fa+6+1dx4obn5UnNvTfNoXYiyHJVlGF6NGlkpRVEUFjDTcD4u+NzJ2e/nNP2Sywy6idAcY9n/GsNZCE4mNt8uoLEodxpKoYKAQfcKz+mR9l1QKaChVJNnlYOUC2kzya7r0jw85KLdJo5j+zLNMtI8xw8jvJFPnKQ4bp9MKXqOa/dhFFFWFWVZkmWZzbOARVniBwFhGFJpjYoiBv0BsefRa1/idDo43S6O606W4zD0PAvY7w+oqmome7XkPKdzdYXf7eG2L2mdnCLT+mqtZ/Lq3Plno35hLpol0x6IwzGqyK1sw0QvJS5HfdZgVYiwP4ZnvzTr3Yl7IIvMZM7deYamB03tkiSxhkTGFBGGSrM10PxONJex5puvSarVIPPRMEBb29t83Njg0+YmO7u7tsj11ESl8Py44vF+xQdHz5j+S7qtoXG40+nY9jEu2loaA6ZMlIBfGOnCWThZBngV09U1XPp3M1ZIZWf36feKJ7uak/C6svN3F11eqs9wOKTValnWZjRN9DKNk04YmnKZpp43q7GKWQ1oDDMTcN5qkabpwhdpWclKyfWhYesHPkEQWNZhGFm2nufZkTTXTGeYfT1yfzFcLnCUKkZhRD6dkFQpgjAkVYXlF2fKznmi8v+bMgMFqhUS633dQGaVZqLSbBFwnuG4EE4SzVmicZXQz4XDSHMaa3vuKGHH1xyMNXuBppqbnD++xrr4PcAGMAAAAABJRU5ErkJggg=="},"images":{"fallback":{"src":"/static/2227958d85bef3ecbfe974892a44ceee/745a1/stft.png","srcSet":"/static/2227958d85bef3ecbfe974892a44ceee/c0e46/stft.png 750w,\n/static/2227958d85bef3ecbfe974892a44ceee/745a1/stft.png 850w","sizes":"100vw"},"sources":[{"srcSet":"/static/2227958d85bef3ecbfe974892a44ceee/5d2dd/stft.webp 750w,\n/static/2227958d85bef3ecbfe974892a44ceee/cf006/stft.webp 850w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.968235294117647}}},"author":[{"id":"Soohwan Kim","bio":"Co-founder/A.I. engineer at TUNiB.","avatar":{"childImageSharp":{"gatsbyImageData":{"layout":"fullWidth","placeholder":{"fallback":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAOCAYAAAAvxDzwAAAACXBIWXMAABYlAAAWJQFJUiTwAAADjElEQVQ4y23Oy48TdQDA8TExEaTttvSBfdHddru7bWc6bWd+8+i0u23n0SdlBRQXZGHDgiiLDwgqXjCIJEaC8cLBmHgwUU9EYzRqQuLJi/4NJB6NF+8evgbOHr7XT76SM5ix7s4QnRHnnm9yZzvBTIswN8JcP5bC67ZoWi7CdtFsn5bloZo+dSNAMYbIIqCmD6mJIVXdR3J6I7r9IQ3T59pZlW+uxzndCbPdC/HZXoq3TikowsfqBOjOmFZ7SrM9p2kfRbXn1O0jKNbsSbJ5BMnu9Gl3+qiiz4OPSvz0QRLHFGwOZH6+neLHT5YRbRerM0J0JrTas/8FZXNKzZgg2es+ZtvDdbs8+iHJl+9lWVP7OG2br26k+PvXJCfmFqoYIpwhTXtMwxpTN0coxhjFmCKbE2rGmKoYPQY9qi2fyzsq//4e5sNXi5TlLqbd5t6lPH/9kuDTG2usqgGi7dO0hzSsIXUzQDaGKGKMLEZUxZCKCJB02+PUpMH9vQwP7ybYnSxyfNrjwumA3dEiD94v8cfnK7xzWsbpujQtH9X0qBs+iuEjC5+a7lPVPSqai9SyelyYFbjkh5hqC3TkKOutLON2nkYpgtuMcT6IszdPslo8SLlcpmUH1A0XRQyQhUtNuFQ1l4o2QFI1h3OTVXbcg+grIfSVBRqlEGopgrYSRi0e4KiIse1mScSeopA/iLADmoZL3Rgg631q+oCq1qeq95FOzhQefpzn7eMpBmqcoBVj04pzwkky0eNMRYIdL8F3twusLe2nslrE6g5pGX1U8/FhD1nvUW1tUJK7SH9+n+DR1zF2vQSvH0lz88wiV45V2J1WuHqsxLXNAnuzNJdfWmK5/BylUgFhb6DZAxR9g7rpIos+Wttja2uE9M9vYb69k+DqfJErWzLOuoVqdKlpDo4jODms4zlFckuHObx4iHwuRi6fYqmYY201x/LqMrnFDNl8gRvvukhNdT9394rcOlNFMxrUhYFmW9imTEsp0mhWKC2nWSokyWSiFA7HSKcXiEb34fayyLVniUQk8rmn2drKIqlqiNcuVrh2ymQS1Nh8sc/GoMEbl7u8eb5D1yqTzkRJJsOkkiGymQVShyLE4yHWKnEq1RjpTIhmI8r2y3mkyloI38uha1n6bpVXrkw5+4LD/S/2uHnrIopaJhZ9hmQqwqHkAXJP8BDhyD46ZoWdMxOGcx/PNfECk/8AbxXdRjRliPoAAAAASUVORK5CYII="},"images":{"fallback":{"src":"/static/a9e6b445142b247ee4cfa66155398bb2/7cf1f/soohwan.png","srcSet":"/static/a9e6b445142b247ee4cfa66155398bb2/34f77/soohwan.png 750w,\n/static/a9e6b445142b247ee4cfa66155398bb2/a94f6/soohwan.png 1080w,\n/static/a9e6b445142b247ee4cfa66155398bb2/7cf1f/soohwan.png 1148w","sizes":"100vw"},"sources":[{"srcSet":"/static/a9e6b445142b247ee4cfa66155398bb2/38420/soohwan.webp 750w,\n/static/a9e6b445142b247ee4cfa66155398bb2/7470d/soohwan.webp 1080w,\n/static/a9e6b445142b247ee4cfa66155398bb2/b5ef6/soohwan.webp 1148w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.6829268292682927}}}}]},"fields":{"readingTime":{"text":"5 min read"},"layout":"","slug":"/dsp_for_audio/"}},"next":{"excerpt":"AI & Speech Processing: Application-2 본 글은 광운대학교 전자공학과 박호종 교수님의 강의를 듣고 작성되었음을 밝힙니다. Speaker Verification and Identification Verification…","frontmatter":{"title":"Sooftware Speech - AI & Speech Processing: Application-2","tags":["speech"],"date":"2020-04-17T10:00:00.000Z","draft":false,"excerpt":null,"image":{"childImageSharp":{"gatsbyImageData":{"layout":"fullWidth","placeholder":{"fallback":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAMCAYAAABiDJ37AAAACXBIWXMAAAsTAAALEwEAmpwYAAACrklEQVQoz42QXUhTYRjHX7cd3aZOj1Ob5kfRFxJRzPXh3BY1j+ecuaOz8iIo+iASskDqVhGKbrz3xrsCL4IsqBtvFbS88sYgglJJIyTd2irtg35xXmERddELf56HPy+/5/88oqm+mm69jfOnUnQbOg/u9vPw7nX2NQRIdqX4Bqxlv7Ce2yD96W9t+ZukP2/KKrZXlhPc34QWPkrscJD+i6c52LST/ft2c+tmP/L9/Mn/PiGEYFdjPTtqA+xprKO0pMQ2pW5cv5b/uL62RjqdJpvNksvlyKQzZDIZVldX2dz8+hvYEo3Q0homHImgmyYdVhLd1El2WhwJhzGTSYxkB216O0YiQdKyONEW5/jJE5iJBFp7O1Z3itTZHjq6OhFzc3NMTEzwfHaWhYUFhoeHGRoaYmRkhIuXL+XTejweirwehKMg71VWVVG7vQ6fvwK3v2zLX1xclFE3NjaYmpoiGAzS3NxMNBolrmk4ChXK1HJUVUU4HVzpvcr09DS379yhszuFy+2mQFFwKApOdxFiaWkpD5ycnETXdUKhEIZhSNnJ/FWVqH4/5RUqY2NjjI+PMzAwwOjoKA17d+GrVFHrA5RWqX8CZ2ZmJDASiWAahrybvYbL496aLgS9vb2c6emRvX0Cp9ct5Sp2oxR7/g201zVNE83QEQVbQFt2HzpymMHBQQqLvXh9PkrKy/CpKsU+H46iQsTy8rIEfv/+Q95G0zRaW1slON6u5YF2woJCBU9pCX19fRw4dIijx47REg5TsaMWf20Apw18MT9PLvuRd+9WmH3+DMuyiMVisnZ0Wn+sbCewa3VNgOqaGjnUDmAkTOoaGxAuJ+L10grzL1/z6s1bVt5/4NHjx9y7f58nT59y7sIFCVS8njxM3tLllElPxuNEY1GCoRD+bdUIxcUvykIN7KMoMm4AAAAASUVORK5CYII="},"images":{"fallback":{"src":"/static/b3321f8e0e847a673a28ec1ffe7b6c8c/c577e/audio_app2.png","srcSet":"/static/b3321f8e0e847a673a28ec1ffe7b6c8c/ca06a/audio_app2.png 750w,\n/static/b3321f8e0e847a673a28ec1ffe7b6c8c/c577e/audio_app2.png 753w","sizes":"100vw"},"sources":[{"srcSet":"/static/b3321f8e0e847a673a28ec1ffe7b6c8c/53639/audio_app2.webp 750w,\n/static/b3321f8e0e847a673a28ec1ffe7b6c8c/71889/audio_app2.webp 753w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.6108897742363878}}},"author":[{"id":"Soohwan Kim","bio":"Co-founder/A.I. engineer at TUNiB.","avatar":{"childImageSharp":{"gatsbyImageData":{"layout":"fullWidth","placeholder":{"fallback":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAOCAYAAAAvxDzwAAAACXBIWXMAABYlAAAWJQFJUiTwAAADjElEQVQ4y23Oy48TdQDA8TExEaTttvSBfdHddru7bWc6bWd+8+i0u23n0SdlBRQXZGHDgiiLDwgqXjCIJEaC8cLBmHgwUU9EYzRqQuLJi/4NJB6NF+8evgbOHr7XT76SM5ix7s4QnRHnnm9yZzvBTIswN8JcP5bC67ZoWi7CdtFsn5bloZo+dSNAMYbIIqCmD6mJIVXdR3J6I7r9IQ3T59pZlW+uxzndCbPdC/HZXoq3TikowsfqBOjOmFZ7SrM9p2kfRbXn1O0jKNbsSbJ5BMnu9Gl3+qiiz4OPSvz0QRLHFGwOZH6+neLHT5YRbRerM0J0JrTas/8FZXNKzZgg2es+ZtvDdbs8+iHJl+9lWVP7OG2br26k+PvXJCfmFqoYIpwhTXtMwxpTN0coxhjFmCKbE2rGmKoYPQY9qi2fyzsq//4e5sNXi5TlLqbd5t6lPH/9kuDTG2usqgGi7dO0hzSsIXUzQDaGKGKMLEZUxZCKCJB02+PUpMH9vQwP7ybYnSxyfNrjwumA3dEiD94v8cfnK7xzWsbpujQtH9X0qBs+iuEjC5+a7lPVPSqai9SyelyYFbjkh5hqC3TkKOutLON2nkYpgtuMcT6IszdPslo8SLlcpmUH1A0XRQyQhUtNuFQ1l4o2QFI1h3OTVXbcg+grIfSVBRqlEGopgrYSRi0e4KiIse1mScSeopA/iLADmoZL3Rgg631q+oCq1qeq95FOzhQefpzn7eMpBmqcoBVj04pzwkky0eNMRYIdL8F3twusLe2nslrE6g5pGX1U8/FhD1nvUW1tUJK7SH9+n+DR1zF2vQSvH0lz88wiV45V2J1WuHqsxLXNAnuzNJdfWmK5/BylUgFhb6DZAxR9g7rpIos+Wttja2uE9M9vYb69k+DqfJErWzLOuoVqdKlpDo4jODms4zlFckuHObx4iHwuRi6fYqmYY201x/LqMrnFDNl8gRvvukhNdT9394rcOlNFMxrUhYFmW9imTEsp0mhWKC2nWSokyWSiFA7HSKcXiEb34fayyLVniUQk8rmn2drKIqlqiNcuVrh2ymQS1Nh8sc/GoMEbl7u8eb5D1yqTzkRJJsOkkiGymQVShyLE4yHWKnEq1RjpTIhmI8r2y3mkyloI38uha1n6bpVXrkw5+4LD/S/2uHnrIopaJhZ9hmQqwqHkAXJP8BDhyD46ZoWdMxOGcx/PNfECk/8AbxXdRjRliPoAAAAASUVORK5CYII="},"images":{"fallback":{"src":"/static/a9e6b445142b247ee4cfa66155398bb2/7cf1f/soohwan.png","srcSet":"/static/a9e6b445142b247ee4cfa66155398bb2/34f77/soohwan.png 750w,\n/static/a9e6b445142b247ee4cfa66155398bb2/a94f6/soohwan.png 1080w,\n/static/a9e6b445142b247ee4cfa66155398bb2/7cf1f/soohwan.png 1148w","sizes":"100vw"},"sources":[{"srcSet":"/static/a9e6b445142b247ee4cfa66155398bb2/38420/soohwan.webp 750w,\n/static/a9e6b445142b247ee4cfa66155398bb2/7470d/soohwan.webp 1080w,\n/static/a9e6b445142b247ee4cfa66155398bb2/b5ef6/soohwan.webp 1148w","type":"image/webp","sizes":"100vw"}]},"width":1,"height":0.6829268292682927}}}}]},"fields":{"readingTime":{"text":"4 min read"},"layout":"","slug":"/audio_app2/"}},"primaryTag":"speech"}},
    "staticQueryHashes": ["3170763342","3229353822"]}