Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Showcase] CallRec — 手机通话录音的 100% 离线档案库(SenseVoice + CAM++ 声纹 + Ollama 摘要 + 关系图谱)

Open Beginner friendly
#3,730 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
1/5
Estimated time
Under an hour
Newbie friendliness
88/100
Issue type
Documentation
Clarity
Clearly specified
Activity status
Active
Tech stack
fastapi, ollama, python, pytorch, sqlite
Domain
documentation

Research direction

Start with docs/community_projects.md and review its existing showcase format. Add CallRec with the repository link, offline workflow, and relevant FunASR integrations from the issue, then verify that the entry meets the project’s showcase conventions and includes any requested screenshots, GIF, or Colab link.

Written by the indexing model from the issue text.

Description

这是什么 / What

CallRec 是一个面向「手机自动通话录音」垂直场景的全本地闭环应用:把成千上万通躺在文件夹里的 .m4a,变成可全文检索、可按人/按时间统计、可看关系图谱、可导出音色克隆的私人档案库。全程离线,音频与转写永不离开本机。

仓库:https://github.com/RevolutionLA/call-recording-archive (MIT)

用到的 FunASR 能力

环节 模型
VAD fsmn-vad
转写 SenseVoiceSmall(language=auto, use_itn)
标点恢复 ct-punc
声纹/说话人分离 CAM++(192 维 embedding + scipy 层次聚类,通话场景强制 k=2)

完整流水线

scan(文件名解析姓名/号码/时间)
→ run(16k 规范化 → VAD → SenseVoice → 标点 → CAM++ → 双人聚类)
→ align(跨通话质心聚类自动锁定「我」→ 联系人归并)
→ summarize(本地 Ollama:摘要/待办/事件/情绪)
→ graph + web 驾驶舱(FastAPI + 自绘 Canvas 力导向图 + ECharts,零 CDN)
→ voices(导出 6–18s 干净人声片段给 Qwen3-TTS 克隆音色)

亮点

  • 通话专属:利用「每通电话你必然在场」的先验自动确定「我」,这是会议/播客工具没有的问题设定;电话 8k 音质下的声纹阈值与聚类策略已实测调优。
  • 断点续跑:每通独立状态,千通级批处理一夜跑完,单通失败不拖垮整批。
  • SQLite 单文件即数据库(WAL 并发),备份=复制一个文件。
  • 零外网依赖:模型走 ModelScope 国内缓存,前端字体与图表库全部本地 vendor。

欢迎 FunASR 用户试用与反馈;如有兴趣,也很乐意被收录进 docs/community_projects.md / use-case showcase。愿意按你们的收录规范补充任何材料(演示截图、GIF、Colab 链接等)。

English summary

CallRec turns folders of phone call recordings into a fully offline personal archive: SenseVoice + fsmn-vad + ct-punc transcription, CAM++ voiceprint diarization that auto-identifies "me" across calls, contact merging, local LLM (Ollama) summaries, full-text search, a hand-drawn Canvas force-directed relationship graph, and clean voice-clip export for Qwen3-TTS cloning. No audio or transcript ever leaves the machine. MIT licensed — feedback and contributions welcome.

Dominant language
Python
Stars
20.5k
Forks
2.1k
Avg merge
15h 21m
Merged PRs (30d)
125

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from modelscope/FunASR

All issues in modelscope/FunASR

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.