CampPlusEmbedder: embeddings differ from reference CAM++ (fbank not mean-normalized, Hamming window, pooling divisor)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 52/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- swift
- Ambito
- audio-video-rtc, machine-learning
Direzione di ricerca
Start with CampPlusEmbedder.embed and the CamPlusPreprocessor and CamPlusPlus model paths, then run the supplied comparison script against the reference CAM++ pipeline. Check the fbank normalization and window settings, followed by segment pooling behavior for partial segments. Done means the shipped embeddings match the reference across the reported clips and pooling lengths.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
CampPlusEmbedder (0.17.5, same on main) doesn't produce the embeddings of the CAM++ model it was converted from. I compared it with the PyTorch reference: campplus_cn_common.bin from iic/speech_campplus_sv_zh-cn_16k-common in 3D-Speaker's CAMPPlus, features as in FunASR's extract_feature. On four 16 kHz clips the shipped pipeline reaches a cosine of 0.35 to 0.59 against the reference embedding of the same clip.
Three causes, largest first:
1. The fbank isn't mean-normalized. FunASR subtracts the per-utterance mean from the fbank (feature - feature.mean(dim=0, keepdim=True) in funasr/models/campplus/utils.py), and so does 3D-Speaker (FBank(mean_nor=True)). Neither CoreML model does it: CamPlusPreprocessor ends at log, and CamPlusPlus casts and transposes feats straight into the first conv. The model card says "CAM++ normalizes the fbank internally", but this graph doesn't. Subtracting the mean in embed raises the cosine to 0.989–0.993.
2. Hamming window. The preprocessor's window constant is Hamming (0.08 at the edges). Kaldi.fbank defaults to Povey, which is what FunASR uses. The reference fbank (Povey, mean-normalized) run through CamPlusPlus gives 0.9975–0.9998, so the window costs roughly another 0.005–0.01.
3. Segment pooling divisor. seg_pooling in CAM++ is F.avg_pool1d(kernel_size=100, stride=100, ceil_mode=True), and PyTorch divides the last, partial segment by the frames it actually holds. The converted avg_pool ops have ceil_mode = true and exclude_padding_from_average = false, so they divide by 100. Fed identical features, CoreML matches torch only when the frame count after the stride-2 TDNN is a multiple of 100:
| fbank frames | after TDNN | cosine CoreML vs torch |
|---|---|---|
| 400 | 200 | 0.9999 |
| 450 | 225 | 0.9831 |
| 500 | 250 | 0.9757 |
| 600 | 300 | 0.9999 |
| 650 | 325 | 0.9734 |
The missing normalization also pulls different speakers together. Cosines between macOS say voices, shipped / with mean subtraction / torch reference:
| pair | shipped | mean subtracted | reference |
|---|---|---|---|
| Daniel – Thomas | 0.778 | 0.505 | 0.509 |
| Anna – Samantha | 0.836 | 0.741 | 0.713 |
| Samantha – Thomas | 0.479 | 0.399 | 0.366 |
We ran into this in a cross-meeting speaker-recognition benchmark on AMI, where CampPlusEmbedder scored different people as the same speaker.
Fixes, in order of impact: subtract the per-utterance fbank mean (in the preprocessor graph or in embed), use a Povey window, and divide the last pooling segment by its real length. I haven't tested whether exclude_padding_from_average = true does that for the ceil-mode overhang.
Comparison script
Needs torch torchaudio coremltools numpy soundfile, the two .mlmodelc folders from FluidInference/campplus-coreml, campplus_cn_common.bin, and DTDNN.py + layers.py from 3D-Speaker under speakerlab/models/campplus/.
import sys, glob, numpy as np, torch, soundfile as sf, coremltools as ct
import torchaudio.compliance.kaldi as K
sys.path.insert(0, '.')
from speakerlab.models.campplus.DTDNN import CAMPPlus
torch.set_grad_enabled(False)
ref = CAMPPlus(feat_dim=80, embedding_size=192)
ref.load_state_dict(torch.load('campplus_cn_common.bin', map_location='cpu')); ref.eval()
pre = ct.models.CompiledMLModel('CamPlusPreprocessor.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
net = ct.models.CompiledMLModel('CamPlusPlus.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
cos = lambda a, b: float(np.dot(a, b) / np.linalg.norm(a) / np.linalg.norm(b))
def ref_feats(w, window='povey'):
f = K.fbank(torch.from_numpy(w)[None], num_mel_bins=80, sample_frequency=16000, dither=0.0, window_type=window)
return f - f.mean(0, keepdim=True)
ref_emb = lambda f: ref(f[None])[0].numpy()
cml_emb = lambda f: net.predict({'feats': np.asarray(f, np.float32)[None]})['embedding'].reshape(-1)
fluid_feats = lambda w: pre.predict({'waveform': (w * 32768).astype(np.float32)[None]})['features'][0] # as CampPlusEmbedder
for path in sorted(glob.glob('audio/*.wav')): # 16 kHz mono
w, _ = sf.read(path, dtype='float32')
r, ff = ref_emb(ref_feats(w)), fluid_feats(w)
print(path,
'shipped', round(cos(cml_emb(ff), r), 4),
'mean subtracted', round(cos(cml_emb(ff - ff.mean(0, keepdims=True)), r), 4),
'reference fbank', round(cos(cml_emb(ref_feats(w).numpy()), r), 4))
w, _ = sf.read(sorted(glob.glob('audio/*.wav'))[0], dtype='float32')
f = ref_feats(w)
for T in (400, 450, 500, 600, 650):
g = f[:T] - f[:T].mean(0, keepdim=True)
print(T, round(cos(cml_emb(g.numpy()), ref_emb(g)), 4))
- Lingua principale
- Swift
- Stelle
- 3k
- Fork
- 451
- Merge medio
- 17h 57m
- PR unite (30g)
- 54
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di FluidInference/FluidAudio
-
CtcKeywordSpotter: input padding is not zeroed, so CTC log-probs are often independent of the audioAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
FluidInference/FluidAudio#991 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
FluidInference/FluidAudio#990 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
FluidInference/FluidAudio#979 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Word end time truncated ~0.5s, opening an inter-word gap that is not silenceForse già presa @maboa l’ha presa 7 giorni fa. Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 30/100
FluidInference/FluidAudio#971 · 5 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 15/100
FluidInference/FluidAudio#967 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di FluidInference/FluidAudio
Issue simili
-
kiosk_set_screensaver_mode ignored: mode is not passed through by the notification parserForse già presa @bgoncal l’ha presa oggi. Apertabug ios
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
home-assistant/iOS#6001 ·
I maintainer di solito rispondono entro 1 giorno
-
area: docs area: ios S3: minor
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
manaflow-ai/cmux#18416 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
FilePath.extension setter crashes with non-ASCII charactersForse già presa @crleonard l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
apple/swift-system#404 ·
I maintainer di solito rispondono entro 1 giorno
-
#️⃣ REX and feebacks 🔍 triage 🧑💻 Developer eXperience 🧰 library
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Orange-OpenSource/ouds-ios#1795 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
RodnaPamet/agrent-ios#172 ·
I maintainer di solito rispondono entro 1 giorno