DiarizationLM: Speaker Diarization Post-Processing with Large Language Models
Quan Wang, Yiling Huang, Guanlong Zhao, Evan B. Clark, Wei Xia, Hank Liao
Google (United States)
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
In this paper, we introduce DiarizationLM, a framework to leverage large language models (LLM) to post-process the outputs from a speaker diarization system.In this framework, the outputs of the automatic speech recognition (ASR) and speaker diarization systems are represented as a compact textual format, which is included in the prompt to an optionally finetuned LLM.The outputs of the LLM can be used as the refined diarization results with the desired enhancement.As a post-processing step, this framework can be easily applied to any off-the-shelf ASR and speaker diarization systems without retraining existing components.Our experiments show that a finetuned PaLM 2-S model can reduce the WDER by rel.55.5% on the Fisher telephone conversation dataset, and rel.44.9% on the Callhome English dataset.
逐年被引趋势
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
计算机 / AISpeech Recognition and Synthesis
Natural Language Processing Techniques
参考文献 0
引用本文 23
按被引量排序,此处列出前 3 条