Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

CampPlusEmbedder: embeddings differ from reference CAM++ (fbank not mean-normalized, Hamming window, pooling divisor)

Aperta
#985 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
52/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
swift

Direzione di ricerca

Start with CampPlusEmbedder.embed and the CamPlusPreprocessor and CamPlusPlus model paths, then run the supplied comparison script against the reference CAM++ pipeline. Check the fbank normalization and window settings, followed by segment pooling behavior for partial segments. Done means the shipped embeddings match the reference across the reported clips and pooling lengths.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

CampPlusEmbedder (0.17.5, same on main) doesn't produce the embeddings of the CAM++ model it was converted from. I compared it with the PyTorch reference: campplus_cn_common.bin from iic/speech_campplus_sv_zh-cn_16k-common in 3D-Speaker's CAMPPlus, features as in FunASR's extract_feature. On four 16 kHz clips the shipped pipeline reaches a cosine of 0.35 to 0.59 against the reference embedding of the same clip.

Three causes, largest first:

1. The fbank isn't mean-normalized. FunASR subtracts the per-utterance mean from the fbank (feature - feature.mean(dim=0, keepdim=True) in funasr/models/campplus/utils.py), and so does 3D-Speaker (FBank(mean_nor=True)). Neither CoreML model does it: CamPlusPreprocessor ends at log, and CamPlusPlus casts and transposes feats straight into the first conv. The model card says "CAM++ normalizes the fbank internally", but this graph doesn't. Subtracting the mean in embed raises the cosine to 0.989–0.993.

2. Hamming window. The preprocessor's window constant is Hamming (0.08 at the edges). Kaldi.fbank defaults to Povey, which is what FunASR uses. The reference fbank (Povey, mean-normalized) run through CamPlusPlus gives 0.9975–0.9998, so the window costs roughly another 0.005–0.01.

3. Segment pooling divisor. seg_pooling in CAM++ is F.avg_pool1d(kernel_size=100, stride=100, ceil_mode=True), and PyTorch divides the last, partial segment by the frames it actually holds. The converted avg_pool ops have ceil_mode = true and exclude_padding_from_average = false, so they divide by 100. Fed identical features, CoreML matches torch only when the frame count after the stride-2 TDNN is a multiple of 100:

fbank frames after TDNN cosine CoreML vs torch
400 200 0.9999
450 225 0.9831
500 250 0.9757
600 300 0.9999
650 325 0.9734

The missing normalization also pulls different speakers together. Cosines between macOS say voices, shipped / with mean subtraction / torch reference:

pair shipped mean subtracted reference
Daniel – Thomas 0.778 0.505 0.509
Anna – Samantha 0.836 0.741 0.713
Samantha – Thomas 0.479 0.399 0.366

We ran into this in a cross-meeting speaker-recognition benchmark on AMI, where CampPlusEmbedder scored different people as the same speaker.

Fixes, in order of impact: subtract the per-utterance fbank mean (in the preprocessor graph or in embed), use a Povey window, and divide the last pooling segment by its real length. I haven't tested whether exclude_padding_from_average = true does that for the ceil-mode overhang.

Comparison script

Needs torch torchaudio coremltools numpy soundfile, the two .mlmodelc folders from FluidInference/campplus-coreml, campplus_cn_common.bin, and DTDNN.py + layers.py from 3D-Speaker under speakerlab/models/campplus/.

import sys, glob, numpy as np, torch, soundfile as sf, coremltools as ct
import torchaudio.compliance.kaldi as K
sys.path.insert(0, '.')
from speakerlab.models.campplus.DTDNN import CAMPPlus

torch.set_grad_enabled(False)
ref = CAMPPlus(feat_dim=80, embedding_size=192)
ref.load_state_dict(torch.load('campplus_cn_common.bin', map_location='cpu')); ref.eval()
pre = ct.models.CompiledMLModel('CamPlusPreprocessor.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
net = ct.models.CompiledMLModel('CamPlusPlus.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
cos = lambda a, b: float(np.dot(a, b) / np.linalg.norm(a) / np.linalg.norm(b))

def ref_feats(w, window='povey'):
    f = K.fbank(torch.from_numpy(w)[None], num_mel_bins=80, sample_frequency=16000, dither=0.0, window_type=window)
    return f - f.mean(0, keepdim=True)
ref_emb = lambda f: ref(f[None])[0].numpy()
cml_emb = lambda f: net.predict({'feats': np.asarray(f, np.float32)[None]})['embedding'].reshape(-1)
fluid_feats = lambda w: pre.predict({'waveform': (w * 32768).astype(np.float32)[None]})['features'][0]  # as CampPlusEmbedder

for path in sorted(glob.glob('audio/*.wav')):  # 16 kHz mono
    w, _ = sf.read(path, dtype='float32')
    r, ff = ref_emb(ref_feats(w)), fluid_feats(w)
    print(path,
          'shipped', round(cos(cml_emb(ff), r), 4),
          'mean subtracted', round(cos(cml_emb(ff - ff.mean(0, keepdims=True)), r), 4),
          'reference fbank', round(cos(cml_emb(ref_feats(w).numpy()), r), 4))

w, _ = sf.read(sorted(glob.glob('audio/*.wav'))[0], dtype='float32')
f = ref_feats(w)
for T in (400, 450, 500, 600, 650):
    g = f[:T] - f[:T].mean(0, keepdim=True)
    print(T, round(cos(cml_emb(g.numpy()), ref_emb(g)), 4))
Lingua principale
Swift
Stelle
3k
Fork
451
Merge medio
17h 57m
PR unite (30g)
54

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di FluidInference/FluidAudio

Tutte le issue di FluidInference/FluidAudio

Issue simili

Altre issue su Swift

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.