Skip to content

Fun-ASR-Nano Hugging Face checkpoint is missing 86 CTC tensors required for timestamps and diarization #3496

Description

@oliver-mee

Summary

The current Hugging Face and ModelScope copies of FunAudioLLM/Fun-ASR-Nano-2512 do not contain the same model.pt.

The Hugging Face file is smaller and lacks every ctc_decoder.* and ctc.* tensor. ASR text still works, but the CTC timestamps required by FunASR's native speaker diarization path are unavailable.

Verified on 2026-08-14 with FunASR 1.4.2.

Artefact comparison

source model.pt size state-dict tensors CTC tensors
Hugging Face, commit 272c57b82523ada6fd87095e955f8e29100979ab 1,971,149,431 bytes 1,261 0
ModelScope 2,127,426,538 bytes 1,347 84 ctc_decoder.* + 2 ctc.*

Complete ModelScope checkpoint SHA-256:

81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499

Complete namespace counts:

audio_encoder  914
audio_adaptor   36
llm            311
ctc_decoder     84
ctc              2

Minimal inspection

from collections import Counter
from huggingface_hub import hf_hub_download
import os
import torch

path = hf_hub_download("FunAudioLLM/Fun-ASR-Nano-2512", "model.pt")
state = torch.load(path, map_location="cpu", mmap=True, weights_only=False)

print(os.path.getsize(path))
print(len(state))
print(Counter(key.split(".", 1)[0] for key in state))

Effect

Using the Hugging Face checkpoint through AutoModel produces text, but no Nano CTC timestamps. That in turn prevents the native spk_model="cam++" diarization path from producing valid speaker-attributed segments.

Replacing only model.pt with the complete ModelScope checkpoint restores character timestamps and native diarization. With the complete checkpoint, AutoModel works on both CPU and Apple MPS.

Expected result

Please synchronise the complete ModelScope model.pt to the Hugging Face repository, or otherwise document why the published artefacts intentionally differ.

It may also be useful for the loader to fail closed when a Nano checkpoint has no CTC tensors, rather than allowing text-only inference to make the checkpoint appear complete.

Related: #3208 and #3211.

Activity

  1. added
    bugSomething isn't working
    needs maintainer decisionWaiting for a maintainer or core author decision, roadmap call, or release-scope confirmation
    on Aug 14, 2026
  2. LauraGPT commented on Aug 14, 2026

    @LauraGPT
    Collaborator

    Independent verification confirms this report.

    Artifact evidence

    source exact object bytes state-dict entries CTC tensors
    Hugging Face at 272c57b82523ada6fd87095e955f8e29100979ab 55ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb 1,971,149,431 1,261 0
    ModelScope complete artifact 81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499 2,127,426,538 1,347 84 ctc_decoder.* + 2 ctc.*

    I downloaded the ModelScope file independently with byte ranges, reassembled it to exactly 2,127,426,538 bytes, and matched the server's X-Linked-Etag. The complete namespace counts are 914 audio_encoder, 36 audio_adaptor, 311 llm, 84 ctc_decoder, and 2 ctc.

    Public-package runtime proof

    Using the public PyPI wheel funasr==1.4.2 on CPU:

    • the complete checkpoint loaded with all keys matched;
    • the official Chinese sample returned 开饭时间:早上九点至下午五点。;
    • CTC output contained 15 character timestamps;
    • fsmn-vad + cam++ produced sentence_info with spk: 0.

    The existing loader protection from #3211 already disables incomplete CTC modules rather than running random timestamp weights, so no additional code fix is required for fail-closed behavior. The remaining fix is the model artifact sync.

    Publication state

    A complete atomic HF update was prepared: replace model.pt and add a bilingual model-card integrity contract with the exact size, SHA-256, and tensor counts. No HF content was changed: the current fine-grained token has an empty model-specific permission override, and backup-branch creation, direct LFS upload, create_pr=True preupload, and discussion creation all returned HTTP 403 before mutation.

    This issue should remain open as a model-publication blocker. A FunAudioLLM model owner needs to upload the complete ModelScope model.pt. After that, I will re-download the HF object, verify SHA-256/state dict, rerun the public 1.4.2 CPU timestamp and diarization tests, and close the issue.

  3. LauraGPT commented on Aug 16, 2026

    @LauraGPT
    Collaborator

    Assigned this publication blocker to @pengzhendong and myself so it has an explicit owner path.

    The complete ModelScope artifact remains available and re-verifies at:

    • bytes: 2,127,426,538
    • SHA-256: 81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499
    • state-dict entries: 1,347
    • CTC tensors: 84 ctc_decoder.* + 2 ctc.*

    The current operations token is an administrator at the FunAudioLLM organization level, but its model-specific override for FunAudioLLM/Fun-ASR-Nano-2512 has no content-write permission, so organization access is overridden for this repository. The browser session is not authenticated, so I did not alter token permissions or create a broader credential.

    @pengzhendong, the remaining owner action is either:

    1. replace the Hugging Face model.pt with the complete object above; or
    2. grant this operations token content-write permission only for FunAudioLLM/Fun-ASR-Nano-2512.

    Once either happens, I will perform the atomic model/card update, independently download the published object, verify the hash and 86 CTC tensors, rerun the public FunASR 1.4.2 CPU timestamp + diarization proof, and close this issue.

  4. LauraGPT commented on Aug 30, 2026

    @LauraGPT
    Collaborator

    重新核验于 2026-08-30:HF 主分支仍是 272c57b8,model.pt 仍为 1,971,149,431 bytes / SHA-256 55ae0d2f,未包含 86 个 CTC tensors。现有 token 的该模型 scoped permissions 仍为空。另验证了最小权限 service-account 方案:当前 token 确有 org.serviceAccounts.write,但 HF 在创建任何账户或凭据之前返回 HTTP 402,说明 FunAudioLLM 当前套餐不支持该企业功能;服务器上没有生成新 token 或半成品。这个 issue 继续保持 open / needs maintainer decision,仍需要模型 owner 给该 repo 补 content-write 或直接上传完整对象。

  5. LauraGPT commented on Aug 31, 2026

    @LauraGPT
    Collaborator

    重新尝试发布于 2026-08-31:候选完整权重已再次通过发布前校验(2,127,426,538 bytes,SHA-256 81fec861…a499,1,347 state-dict entries,86 个 CTC tensors),英中模型卡与回滚备份也已准备完成。当前 token 对仓库的 HF auth_check 可以通过,但原子 create_commit 在任何 commit 创建前的 LFS batch 阶段仍返回 HTTP 403;auth_check 只验证仓库访问,不代表 LFS content-write。发布失败后重新读取公开仓库,HEAD 仍为 272c57b8,model.pt 仍为旧的 1,971,149,431 bytes / SHA-256 55ae0d2f,确认远端没有半成品或元数据变更。这个 issue 继续保持开放,等待该模型仓库明确授予 LFS/content write 或由模型 owner 上传完整对象;之后仍需公开回下载、SHA/张量计数及真实时间戳推理复验,不能因本地候选通过而关闭。

  6. LauraGPT commented on Sep 2, 2026

    @LauraGPT
    Collaborator

    重新按当前官方 HF Hub revision 做了键级核验,结论是这个边界目前仍然存在,不能称为已修复:

    • FunAudioLLM/Fun-ASR-Nano-2512-hf 当前 revision e45779dbfd43717e676ec718ecb215b36e848b18 的 model.safetensors header 包含 1,540 个 tensor,其中 CTC tensor 为 0。
    • 正在审阅的 Transformers 转换脚本也明确将 CTC branch 标为“不被 HF generation model 使用”的未转换键。该 PR 的 FunAsrNanoForConditionalGeneration 生成实现没有 CTC 引用。

    所以,这个 HF checkpoint 可以用于该 Transformers 生成路径,但不能被当作原生 Fun-ASR-Nano checkpoint 的完整功能替代品;尤其不能承诺它具备依赖原生 CTC 分支的时间戳/diarization 能力。需要这些能力时,请继续使用官方原生 checkpoint 及 FunASR/vLLM 路径。

    我已同时修正了 HF processor 的 LFR 字段名并请求 Transformers 维护者复核,但那一项只影响特征提取配置,不会补齐 CTC 权重。本 issue 保持打开,后续需要明确决定:是为 HF 发布完整 CTC 支持,还是在模型卡与文档中把这一能力边界写成正式契约。

  7. LauraGPT commented on Sep 2, 2026

    @LauraGPT
    Collaborator

    补充:上述边界现已写入官方 HF 模型卡并完成 Hub 回读验证。当前模型卡 revision 为 f1ff8a9a371b51c01d8dfe9c4d6ff00a45fd4aa6:
    https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-2512-hf

    模型卡明确说明 Transformers checkpoint 是 generation-based transcription 路径,不含 native CTC branch;需要 CTC-dependent timestamps 或 speaker diarization 时,请使用原生 checkpoint 与 FunASR/vLLM 路径。

  8. LauraGPT commented on Sep 2, 2026

    @LauraGPT
    Collaborator

    最新一次发布恢复已实际完成完整对象验证,但 LFS 写权限仍未授予,因此不能声称已经同步:

    • 从官方 ModelScope FunAudioLLM/Fun-ASR-Nano-2512 恢复的 model.pt 为 2,127,426,538 bytes,SHA-256 81fec8616083c69377f3ceef36aba3655660ee0ca69a5d4a1e9810cd340ca499。
    • 直接读取 state dict:1,347 entries,84 个 ctc_decoder.* 加 2 个 ctc.*,共 86 个 CTC tensors。
    • 新的发布前快照已保存;随后对 model.pt、README.md、README_zh.md 发起单次 Hugging Face 原子 create_commit。
    • 提交在创建 commit 之前的 ...git/info/lfs/objects/batch 阶段得到 HTTP 403,未上传 LFS 对象,也未更新任何模型卡。
    • 公开回读仍为 272c57b82523ada6fd87095e955f8e29100979ab,model.pt 仍是 1,971,149,431 bytes / 55ae0d2f...60bb。

    因此阻塞点已缩小为 FunAudioLLM/Fun-ASR-Nano-2512 的 Hugging Face LFS content-write scope;普通 repo 访问与组织成员身份不足。候选完整权重保留在受控路径,issue 继续保持 open。获得该仓库的 LFS 内容写权限后,将再次使用同一原子提交并做公开回下载、SHA/张量计数与真实时间戳推理复验。

  9. qiulang commented on Sep 2, 2026

    @qiulang
    Contributor

    One downstream consequence worth deciding before the sync lands.

    FunAudioLLM/Fun-ASR-Nano-2512-vllm was converted from the current HF object, and both its provenance file and its conversion script pin that exact source:

    # convert_from_official.py
    SOURCE_REVISION  = "272c57b82523ada6fd87095e955f8e29100979ab"
    SOURCE_SHA256    = "55ae0d2fee369f0f11cce0795f6927934ad17cf11b278a7e56a51272074160bb"
    EXPECTED_TENSORS = 1261

    Once model.pt is replaced, the SHA-256 assertion fails. That is correct fail-closed behaviour, but it means anyone re-running the script against the new HEAD gets an error with no indication of which revision to use instead.

    Note also that convert_from_official.py does no filtering — it is torch.load straight into save_file(state_dict) — so re-running it against a 1,347-tensor source would emit 1,347 tensors, changing the artifact SHA and adding 86 CTC keys that have no counterpart in the vLLM module tree.

    The functional side

    The -vllm artifact does not need the CTC tensors. FunASRForConditionalGeneration in vllm/model_executor/models/funasr.py has no CTC module, and hf_to_vllm_mapper only covers audio_encoder.*, audio_adaptor.*, llm.model.* and llm.lm_head. For that architecture 1,261 is already complete.

    Three options

    1. Add an explicit CTC filter, keep 1,261. The non-CTC tensors are unchanged, i.e. the vLLM-native artifact CTC-free by design, so the output may well be byte-identical to today's artifact (safetensors writes no metadata here, though key ordering would need checking). If it is, the artifact SHA stays stable and only the provenance file needs updating.

    2. Re-base to the complete source and pass everything through (1,347). Requires first confirming whether vLLM's AutoWeightsLoader ignores or rejects the unmapped CTC keys.

    3. Change nothing. The script pins a revision, not HEAD, and 272c57b8 remains reachable in history, so the artifact stays self-consistent and verifiable. Only a documentation note is needed to say it intentionally points at a superseded revision.

  10. LauraGPT commented on Sep 2, 2026

    @LauraGPT
    Collaborator

    已按候选完整 checkpoint 做了逐张量对照,决定采用“显式 CTC 过滤、保持 1,261 个 vLLM-required tensors”的契约。

    验证不是只看 module 名称:

    • 完整源 model.pt:1,347 tensors;
    • ctc.* / ctc_decoder.*:恰好 86 tensors;
    • 过滤后:1,261 tensors;
    • 与当前 vLLM 产物来源的 1,261 项逐项比较,key、shape、dtype 与 tensor value 全部相同。

    因此 vLLM package 会继续明确为 CTC-free by design:这些 CTC tensors 没有 vLLM FunASRForConditionalGeneration 的对应模块,直接透传既不能带来 timestamps/diarization 功能,也会把 loader 兼容性变成不必要的不确定项。

    在官方 HF 源的 LFS content-write 权限恢复前,当前 -vllm artifact 仍保持绑定旧 revision,不提前改 provenance 或制造不可回放的新版本。官方完整源真正同步后,我会先将 converter 改为显式断言“raw 1,347 -> exclude only 86 named CTC keys -> retain 1,261”,再重建、校验 artifact SHA、做 vLLM 实际加载/转写,并同步新的 source revision/SHA 到 provenance。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

bugSomething isn't workingneeds maintainer decisionWaiting for a maintainer or core author decision, roadmap call, or release-scope confirmation

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions