CampPlusEmbedder: embeddings differ from reference CAM++ (fbank not mean-normalized, Hamming window, pooling divisor)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- swift
Línea de trabajo
Start with CampPlusEmbedder.embed and the CamPlusPreprocessor and CamPlusPlus model paths, then run the supplied comparison script against the reference CAM++ pipeline. Check the fbank normalization and window settings, followed by segment pooling behavior for partial segments. Done means the shipped embeddings match the reference across the reported clips and pooling lengths.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
CampPlusEmbedder (0.17.5, same on main) doesn't produce the embeddings of the CAM++ model it was converted from. I compared it with the PyTorch reference: campplus_cn_common.bin from iic/speech_campplus_sv_zh-cn_16k-common in 3D-Speaker's CAMPPlus, features as in FunASR's extract_feature. On four 16 kHz clips the shipped pipeline reaches a cosine of 0.35 to 0.59 against the reference embedding of the same clip.
Three causes, largest first:
1. The fbank isn't mean-normalized. FunASR subtracts the per-utterance mean from the fbank (feature - feature.mean(dim=0, keepdim=True) in funasr/models/campplus/utils.py), and so does 3D-Speaker (FBank(mean_nor=True)). Neither CoreML model does it: CamPlusPreprocessor ends at log, and CamPlusPlus casts and transposes feats straight into the first conv. The model card says "CAM++ normalizes the fbank internally", but this graph doesn't. Subtracting the mean in embed raises the cosine to 0.989–0.993.
2. Hamming window. The preprocessor's window constant is Hamming (0.08 at the edges). Kaldi.fbank defaults to Povey, which is what FunASR uses. The reference fbank (Povey, mean-normalized) run through CamPlusPlus gives 0.9975–0.9998, so the window costs roughly another 0.005–0.01.
3. Segment pooling divisor. seg_pooling in CAM++ is F.avg_pool1d(kernel_size=100, stride=100, ceil_mode=True), and PyTorch divides the last, partial segment by the frames it actually holds. The converted avg_pool ops have ceil_mode = true and exclude_padding_from_average = false, so they divide by 100. Fed identical features, CoreML matches torch only when the frame count after the stride-2 TDNN is a multiple of 100:
| fbank frames | after TDNN | cosine CoreML vs torch |
|---|---|---|
| 400 | 200 | 0.9999 |
| 450 | 225 | 0.9831 |
| 500 | 250 | 0.9757 |
| 600 | 300 | 0.9999 |
| 650 | 325 | 0.9734 |
The missing normalization also pulls different speakers together. Cosines between macOS say voices, shipped / with mean subtraction / torch reference:
| pair | shipped | mean subtracted | reference |
|---|---|---|---|
| Daniel – Thomas | 0.778 | 0.505 | 0.509 |
| Anna – Samantha | 0.836 | 0.741 | 0.713 |
| Samantha – Thomas | 0.479 | 0.399 | 0.366 |
We ran into this in a cross-meeting speaker-recognition benchmark on AMI, where CampPlusEmbedder scored different people as the same speaker.
Fixes, in order of impact: subtract the per-utterance fbank mean (in the preprocessor graph or in embed), use a Povey window, and divide the last pooling segment by its real length. I haven't tested whether exclude_padding_from_average = true does that for the ceil-mode overhang.
Comparison script
Needs torch torchaudio coremltools numpy soundfile, the two .mlmodelc folders from FluidInference/campplus-coreml, campplus_cn_common.bin, and DTDNN.py + layers.py from 3D-Speaker under speakerlab/models/campplus/.
import sys, glob, numpy as np, torch, soundfile as sf, coremltools as ct
import torchaudio.compliance.kaldi as K
sys.path.insert(0, '.')
from speakerlab.models.campplus.DTDNN import CAMPPlus
torch.set_grad_enabled(False)
ref = CAMPPlus(feat_dim=80, embedding_size=192)
ref.load_state_dict(torch.load('campplus_cn_common.bin', map_location='cpu')); ref.eval()
pre = ct.models.CompiledMLModel('CamPlusPreprocessor.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
net = ct.models.CompiledMLModel('CamPlusPlus.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
cos = lambda a, b: float(np.dot(a, b) / np.linalg.norm(a) / np.linalg.norm(b))
def ref_feats(w, window='povey'):
f = K.fbank(torch.from_numpy(w)[None], num_mel_bins=80, sample_frequency=16000, dither=0.0, window_type=window)
return f - f.mean(0, keepdim=True)
ref_emb = lambda f: ref(f[None])[0].numpy()
cml_emb = lambda f: net.predict({'feats': np.asarray(f, np.float32)[None]})['embedding'].reshape(-1)
fluid_feats = lambda w: pre.predict({'waveform': (w * 32768).astype(np.float32)[None]})['features'][0] # as CampPlusEmbedder
for path in sorted(glob.glob('audio/*.wav')): # 16 kHz mono
w, _ = sf.read(path, dtype='float32')
r, ff = ref_emb(ref_feats(w)), fluid_feats(w)
print(path,
'shipped', round(cos(cml_emb(ff), r), 4),
'mean subtracted', round(cos(cml_emb(ff - ff.mean(0, keepdims=True)), r), 4),
'reference fbank', round(cos(cml_emb(ref_feats(w).numpy()), r), 4))
w, _ = sf.read(sorted(glob.glob('audio/*.wav'))[0], dtype='float32')
f = ref_feats(w)
for T in (400, 450, 500, 600, 650):
g = f[:T] - f[:T].mean(0, keepdim=True)
print(T, round(cos(cml_emb(g.numpy()), ref_emb(g)), 4))
- Lenguaje dominante
- Swift
- Estrellas
- 3k
- Forks
- 451
- Merge medio
- 17 h 57 min
- PR fusionados (30 d)
- 54
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de FluidInference/FluidAudio
-
CtcKeywordSpotter: input padding is not zeroed, so CTC log-probs are often independent of the audioAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
FluidInference/FluidAudio#991 ·
Los mantenedores suelen responder en 1 día
-
Request TTS and Diarizer opt-out traits: ASR-only consumers ship 1.0 MB of unreachable LuxTTS resources and compile 3.1 MB of unused subsystems (cf. #880/#888)Posiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
FluidInference/FluidAudio#990 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 30/100
FluidInference/FluidAudio#979 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Word end time truncated ~0.5s, opening an inter-word gap that is not silencePosiblemente ocupada @maboa la tomó hace 8 días. Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 30/100
FluidInference/FluidAudio#971 · 5 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 15/100
FluidInference/FluidAudio#967 · 1 comentario · 1 reacción ·
Los mantenedores suelen responder en 1 día
Todos los issues de FluidInference/FluidAudio
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 77/100
Los mantenedores suelen responder en 1 día
-
bug milestone-qa mobile
Dificultad 2/5 1-3 horas Aptitud para principiantes 90/100
lognorman20/monaco#4011 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 67/100
Cocoanetics/SwiftScript#22 ·
-
🚀 [firebase_core] Bump Firebase iOS SDK (12.19.0 → 13.0.0)Posiblemente ocupada @SelaseKay la tomó hoy. AbiertoNeeds Attention type: enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
firebase/flutterfire#18769 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
🎉 Add NijiaAbiertoaddition
Dificultad 1/5 1-3 horas Aptitud para principiantes 68/100
jaywcjlove/awesome-mac#3269 ·
Los mantenedores suelen responder en 1 día