Repository navigation
Fun-ASR-Nano Hugging Face checkpoint is missing 86 CTC tensors required for timestamps and diarization #3496
Description
Activity
- addedbugSomething isn't workingSomething isn't workingneeds maintainer decisionWaiting for a maintainer or core author decision, roadmap call, or release-scope confirmationWaiting for a maintainer or core author decision, roadmap call, or release-scope confirmation
on Aug 14, 2026 Independent verification confirms this report.
Artifact evidence
source exact object bytes state-dict entries CTC tensors Hugging Face at 272c57b82523ada6fd87095e955f8e29100979ab55ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb1,971,149,431 1,261 0 ModelScope complete artifact 81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca4992,127,426,538 1,347 84 ctc_decoder.*+ 2ctc.*I downloaded the ModelScope file independently with byte ranges, reassembled it to exactly 2,127,426,538 bytes, and matched the server's
X-Linked-Etag. The complete namespace counts are 914audio_encoder, 36audio_adaptor, 311llm, 84ctc_decoder, and 2ctc.Public-package runtime proof
Using the public PyPI wheel
funasr==1.4.2on CPU:- the complete checkpoint loaded with all keys matched;
- the official Chinese sample returned
开饭时间:早上九点至下午五点。; - CTC output contained 15 character timestamps;
fsmn-vad + cam++producedsentence_infowithspk: 0.
The existing loader protection from #3211 already disables incomplete CTC modules rather than running random timestamp weights, so no additional code fix is required for fail-closed behavior. The remaining fix is the model artifact sync.
Publication state
A complete atomic HF update was prepared: replace
model.ptand add a bilingual model-card integrity contract with the exact size, SHA-256, and tensor counts. No HF content was changed: the current fine-grained token has an empty model-specific permission override, and backup-branch creation, direct LFS upload,create_pr=Truepreupload, and discussion creation all returned HTTP 403 before mutation.This issue should remain open as a model-publication blocker. A FunAudioLLM model owner needs to upload the complete ModelScope
model.pt. After that, I will re-download the HF object, verify SHA-256/state dict, rerun the public 1.4.2 CPU timestamp and diarization tests, and close the issue.Assigned this publication blocker to @pengzhendong and myself so it has an explicit owner path.
The complete ModelScope artifact remains available and re-verifies at:
- bytes:
2,127,426,538 - SHA-256:
81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499 - state-dict entries:
1,347 - CTC tensors:
84 ctc_decoder.* + 2 ctc.*
The current operations token is an administrator at the FunAudioLLM organization level, but its model-specific override for
FunAudioLLM/Fun-ASR-Nano-2512has no content-write permission, so organization access is overridden for this repository. The browser session is not authenticated, so I did not alter token permissions or create a broader credential.@pengzhendong, the remaining owner action is either:
- replace the Hugging Face
model.ptwith the complete object above; or - grant this operations token content-write permission only for
FunAudioLLM/Fun-ASR-Nano-2512.
Once either happens, I will perform the atomic model/card update, independently download the published object, verify the hash and 86 CTC tensors, rerun the public FunASR 1.4.2 CPU timestamp + diarization proof, and close this issue.
- bytes:
重新核验于 2026-08-30:HF 主分支仍是 272c57b8,model.pt 仍为 1,971,149,431 bytes / SHA-256 55ae0d2f,未包含 86 个 CTC tensors。现有 token 的该模型 scoped permissions 仍为空。另验证了最小权限 service-account 方案:当前 token 确有 org.serviceAccounts.write,但 HF 在创建任何账户或凭据之前返回 HTTP 402,说明 FunAudioLLM 当前套餐不支持该企业功能;服务器上没有生成新 token 或半成品。这个 issue 继续保持 open / needs maintainer decision,仍需要模型 owner 给该 repo 补 content-write 或直接上传完整对象。
重新尝试发布于 2026-08-31:候选完整权重已再次通过发布前校验(2,127,426,538 bytes,SHA-256 81fec861…a499,1,347 state-dict entries,86 个 CTC tensors),英中模型卡与回滚备份也已准备完成。当前 token 对仓库的 HF auth_check 可以通过,但原子 create_commit 在任何 commit 创建前的 LFS batch 阶段仍返回 HTTP 403;auth_check 只验证仓库访问,不代表 LFS content-write。发布失败后重新读取公开仓库,HEAD 仍为 272c57b8,model.pt 仍为旧的 1,971,149,431 bytes / SHA-256 55ae0d2f,确认远端没有半成品或元数据变更。这个 issue 继续保持开放,等待该模型仓库明确授予 LFS/content write 或由模型 owner 上传完整对象;之后仍需公开回下载、SHA/张量计数及真实时间戳推理复验,不能因本地候选通过而关闭。
重新按当前官方 HF Hub revision 做了键级核验,结论是这个边界目前仍然存在,不能称为已修复:
- FunAudioLLM/Fun-ASR-Nano-2512-hf 当前 revision e45779dbfd43717e676ec718ecb215b36e848b18 的 model.safetensors header 包含 1,540 个 tensor,其中 CTC tensor 为 0。
- 正在审阅的 Transformers 转换脚本也明确将 CTC branch 标为“不被 HF generation model 使用”的未转换键。该 PR 的 FunAsrNanoForConditionalGeneration 生成实现没有 CTC 引用。
所以,这个 HF checkpoint 可以用于该 Transformers 生成路径,但不能被当作原生 Fun-ASR-Nano checkpoint 的完整功能替代品;尤其不能承诺它具备依赖原生 CTC 分支的时间戳/diarization 能力。需要这些能力时,请继续使用官方原生 checkpoint 及 FunASR/vLLM 路径。
我已同时修正了 HF processor 的 LFR 字段名并请求 Transformers 维护者复核,但那一项只影响特征提取配置,不会补齐 CTC 权重。本 issue 保持打开,后续需要明确决定:是为 HF 发布完整 CTC 支持,还是在模型卡与文档中把这一能力边界写成正式契约。
补充:上述边界现已写入官方 HF 模型卡并完成 Hub 回读验证。当前模型卡 revision 为 f1ff8a9a371b51c01d8dfe9c4d6ff00a45fd4aa6:
https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf模型卡明确说明 Transformers checkpoint 是 generation-based transcription 路径,不含 native CTC branch;需要 CTC-dependent timestamps 或 speaker diarization 时,请使用原生 checkpoint 与 FunASR/vLLM 路径。
最新一次发布恢复已实际完成完整对象验证,但 LFS 写权限仍未授予,因此不能声称已经同步:
- 从官方 ModelScope
FunAudioLLM/Fun-ASR-Nano-2512恢复的model.pt为 2,127,426,538 bytes,SHA-25681fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499。 - 直接读取 state dict:1,347 entries,84 个
ctc_decoder.*加 2 个ctc.*,共 86 个 CTC tensors。 - 新的发布前快照已保存;随后对
model.pt、README.md、README_zh.md发起单次 Hugging Face 原子create_commit。 - 提交在创建 commit 之前的
...git/info/lfs/objects/batch阶段得到 HTTP 403,未上传 LFS 对象,也未更新任何模型卡。 - 公开回读仍为
272c57b82523ada6fd87095e955f8e29100979ab,model.pt仍是 1,971,149,431 bytes /55ae0d2f...60bb。
因此阻塞点已缩小为
FunAudioLLM/Fun-ASR-Nano-2512的 Hugging Face LFS content-write scope;普通 repo 访问与组织成员身份不足。候选完整权重保留在受控路径,issue 继续保持 open。获得该仓库的 LFS 内容写权限后,将再次使用同一原子提交并做公开回下载、SHA/张量计数与真实时间戳推理复验。- 从官方 ModelScope
One downstream consequence worth deciding before the sync lands.
FunAudioLLM/Fun-ASR-Nano-2512-vllmwas converted from the current HF object, and both its provenance file and its conversion script pin that exact source:# convert_from_official.py SOURCE_REVISION = "272c57b82523ada6fd87095e955f8e29100979ab" SOURCE_SHA256 = "55ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb" EXPECTED_TENSORS = 1261
Once
model.ptis replaced, the SHA-256 assertion fails. That is correct fail-closed behaviour, but it means anyone re-running the script against the new HEAD gets an error with no indication of which revision to use instead.Note also that
convert_from_official.pydoes no filtering — it istorch.loadstraight intosave_file(state_dict)— so re-running it against a 1,347-tensor source would emit 1,347 tensors, changing the artifact SHA and adding 86 CTC keys that have no counterpart in the vLLM module tree.The functional side
The
-vllmartifact does not need the CTC tensors.FunASRForConditionalGenerationinvllm/model_executor/models/funasr.pyhas no CTC module, andhf_to_vllm_mapperonly coversaudio_encoder.*,audio_adaptor.*,llm.model.*andllm.lm_head. For that architecture 1,261 is already complete.Three options
-
Add an explicit CTC filter, keep 1,261. The non-CTC tensors are unchanged, i.e. the vLLM-native artifact CTC-free by design, so the output may well be byte-identical to today's artifact (safetensors writes no metadata here, though key ordering would need checking). If it is, the artifact SHA stays stable and only the provenance file needs updating.
-
Re-base to the complete source and pass everything through (1,347). Requires first confirming whether vLLM's
AutoWeightsLoaderignores or rejects the unmapped CTC keys. -
Change nothing. The script pins a revision, not HEAD, and
272c57b8remains reachable in history, so the artifact stays self-consistent and verifiable. Only a documentation note is needed to say it intentionally points at a superseded revision.
-
已按候选完整 checkpoint 做了逐张量对照,决定采用“显式 CTC 过滤、保持 1,261 个 vLLM-required tensors”的契约。
验证不是只看 module 名称:
- 完整源
model.pt:1,347 tensors; ctc.*/ctc_decoder.*:恰好 86 tensors;- 过滤后:1,261 tensors;
- 与当前 vLLM 产物来源的 1,261 项逐项比较,key、shape、dtype 与 tensor value 全部相同。
因此 vLLM package 会继续明确为 CTC-free by design:这些 CTC tensors 没有 vLLM
FunASRForConditionalGeneration的对应模块,直接透传既不能带来 timestamps/diarization 功能,也会把 loader 兼容性变成不必要的不确定项。在官方 HF 源的 LFS content-write 权限恢复前,当前
-vllmartifact 仍保持绑定旧 revision,不提前改 provenance 或制造不可回放的新版本。官方完整源真正同步后,我会先将 converter 改为显式断言“raw 1,347 -> exclude only 86 named CTC keys -> retain 1,261”,再重建、校验 artifact SHA、做 vLLM 实际加载/转写,并同步新的 source revision/SHA 到 provenance。- 完整源
Summary
The current Hugging Face and ModelScope copies of
FunAudioLLM/Fun-ASR-Nano-2512do not contain the samemodel.pt.The Hugging Face file is smaller and lacks every
ctc_decoder.*andctc.*tensor. ASR text still works, but the CTC timestamps required by FunASR's native speaker diarization path are unavailable.Verified on 2026-08-14 with FunASR 1.4.2.
Artefact comparison
272c57b82523ada6fd87095e955f8e29100979abctc_decoder.*+ 2ctc.*Complete ModelScope checkpoint SHA-256:
81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499Complete namespace counts:
Minimal inspection
Effect
Using the Hugging Face checkpoint through
AutoModelproduces text, but no Nano CTC timestamps. That in turn prevents the nativespk_model="cam++"diarization path from producing valid speaker-attributed segments.Replacing only
model.ptwith the complete ModelScope checkpoint restores character timestamps and native diarization. With the complete checkpoint,AutoModelworks on both CPU and Apple MPS.Expected result
Please synchronise the complete ModelScope
model.ptto the Hugging Face repository, or otherwise document why the published artefacts intentionally differ.It may also be useful for the loader to fail closed when a Nano checkpoint has no CTC tensors, rather than allowing text-only inference to make the checkpoint appear complete.
Related: #3208 and #3211.