videoaudioASR_mapper

July 17, 2026 · View on GitHub

Perform automatic speech recognition (ASR) on video audio streams using the SenseVoiceSmall model. This operator extracts audio from videos and transcribes speech content. It must be operated after video_tagging_from_audio_mapper, as it only processes videos tagged as containing speech.

使用 SenseVoiceSmall 模型对视频音频流进行自动语音识别(ASR)。此算子从视频中提取音频并转录语音内容。它必须在 video_tagging_from_audio_mapper 之后运行,因为仅处理被标记为包含语音的视频。

Type 算子类型: mapper

Tags 标签: gpu, hf, video

🔧 Parameter Configuration 参数配置

name 参数名type 类型default 默认值desc 说明
model_dir_ASR<class 'str'>'FunAudioLLM/SenseVoiceSmall'path to the SenseVoiceSmall ASR model.
speech_ASR<class 'str'>'speech_ASR'field name to store the ASR results.
args''extra args
kwargs''extra args