Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

CampPlusEmbedder: embeddings differ from reference CAM++ (fbank not mean-normalized, Hamming window, pooling divisor)

Abierto
#985 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
52/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
swift

Línea de trabajo

Start with CampPlusEmbedder.embed and the CamPlusPreprocessor and CamPlusPlus model paths, then run the supplied comparison script against the reference CAM++ pipeline. Check the fbank normalization and window settings, followed by segment pooling behavior for partial segments. Done means the shipped embeddings match the reference across the reported clips and pooling lengths.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

CampPlusEmbedder (0.17.5, same on main) doesn't produce the embeddings of the CAM++ model it was converted from. I compared it with the PyTorch reference: campplus_cn_common.bin from iic/speech_campplus_sv_zh-cn_16k-common in 3D-Speaker's CAMPPlus, features as in FunASR's extract_feature. On four 16 kHz clips the shipped pipeline reaches a cosine of 0.35 to 0.59 against the reference embedding of the same clip.

Three causes, largest first:

1. The fbank isn't mean-normalized. FunASR subtracts the per-utterance mean from the fbank (feature - feature.mean(dim=0, keepdim=True) in funasr/models/campplus/utils.py), and so does 3D-Speaker (FBank(mean_nor=True)). Neither CoreML model does it: CamPlusPreprocessor ends at log, and CamPlusPlus casts and transposes feats straight into the first conv. The model card says "CAM++ normalizes the fbank internally", but this graph doesn't. Subtracting the mean in embed raises the cosine to 0.989–0.993.

2. Hamming window. The preprocessor's window constant is Hamming (0.08 at the edges). Kaldi.fbank defaults to Povey, which is what FunASR uses. The reference fbank (Povey, mean-normalized) run through CamPlusPlus gives 0.9975–0.9998, so the window costs roughly another 0.005–0.01.

3. Segment pooling divisor. seg_pooling in CAM++ is F.avg_pool1d(kernel_size=100, stride=100, ceil_mode=True), and PyTorch divides the last, partial segment by the frames it actually holds. The converted avg_pool ops have ceil_mode = true and exclude_padding_from_average = false, so they divide by 100. Fed identical features, CoreML matches torch only when the frame count after the stride-2 TDNN is a multiple of 100:

fbank frames after TDNN cosine CoreML vs torch
400 200 0.9999
450 225 0.9831
500 250 0.9757
600 300 0.9999
650 325 0.9734

The missing normalization also pulls different speakers together. Cosines between macOS say voices, shipped / with mean subtraction / torch reference:

pair shipped mean subtracted reference
Daniel – Thomas 0.778 0.505 0.509
Anna – Samantha 0.836 0.741 0.713
Samantha – Thomas 0.479 0.399 0.366

We ran into this in a cross-meeting speaker-recognition benchmark on AMI, where CampPlusEmbedder scored different people as the same speaker.

Fixes, in order of impact: subtract the per-utterance fbank mean (in the preprocessor graph or in embed), use a Povey window, and divide the last pooling segment by its real length. I haven't tested whether exclude_padding_from_average = true does that for the ceil-mode overhang.

Comparison script

Needs torch torchaudio coremltools numpy soundfile, the two .mlmodelc folders from FluidInference/campplus-coreml, campplus_cn_common.bin, and DTDNN.py + layers.py from 3D-Speaker under speakerlab/models/campplus/.

import sys, glob, numpy as np, torch, soundfile as sf, coremltools as ct
import torchaudio.compliance.kaldi as K
sys.path.insert(0, '.')
from speakerlab.models.campplus.DTDNN import CAMPPlus

torch.set_grad_enabled(False)
ref = CAMPPlus(feat_dim=80, embedding_size=192)
ref.load_state_dict(torch.load('campplus_cn_common.bin', map_location='cpu')); ref.eval()
pre = ct.models.CompiledMLModel('CamPlusPreprocessor.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
net = ct.models.CompiledMLModel('CamPlusPlus.mlmodelc', compute_units=ct.ComputeUnit.CPU_ONLY)
cos = lambda a, b: float(np.dot(a, b) / np.linalg.norm(a) / np.linalg.norm(b))

def ref_feats(w, window='povey'):
    f = K.fbank(torch.from_numpy(w)[None], num_mel_bins=80, sample_frequency=16000, dither=0.0, window_type=window)
    return f - f.mean(0, keepdim=True)
ref_emb = lambda f: ref(f[None])[0].numpy()
cml_emb = lambda f: net.predict({'feats': np.asarray(f, np.float32)[None]})['embedding'].reshape(-1)
fluid_feats = lambda w: pre.predict({'waveform': (w * 32768).astype(np.float32)[None]})['features'][0]  # as CampPlusEmbedder

for path in sorted(glob.glob('audio/*.wav')):  # 16 kHz mono
    w, _ = sf.read(path, dtype='float32')
    r, ff = ref_emb(ref_feats(w)), fluid_feats(w)
    print(path,
          'shipped', round(cos(cml_emb(ff), r), 4),
          'mean subtracted', round(cos(cml_emb(ff - ff.mean(0, keepdims=True)), r), 4),
          'reference fbank', round(cos(cml_emb(ref_feats(w).numpy()), r), 4))

w, _ = sf.read(sorted(glob.glob('audio/*.wav'))[0], dtype='float32')
f = ref_feats(w)
for T in (400, 450, 500, 600, 650):
    g = f[:T] - f[:T].mean(0, keepdim=True)
    print(T, round(cos(cml_emb(g.numpy()), ref_emb(g)), 4))
Lenguaje dominante
Swift
Estrellas
3k
Forks
451
Merge medio
17 h 57 min
PR fusionados (30 d)
54

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de FluidInference/FluidAudio

Todos los issues de FluidInference/FluidAudio

Issues similares

Más issues de Swift

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.