MaixCAM2: IVPS pool blocks leak on process exit (even with cam.close()), next camera process hangs in __AX_OSAL_SYNC_wait_interruptible
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Error
- Claridad
- Necesita aclaración
- Estado de actividad
- Activo
- Área
- computer-vision, embedded-iot
Línea de trabajo
Start with components/vision/port/maixcam2/maix_camera_maixcam2.cpp, focusing on Camera::close(), channel teardown, and frame handling. Reproduce repeated process exits while monitoring /proc/ax_proc/pool; done means identifying a supported cleanup sequence or confirming that the binary driver fails to reclaim IVPS blocks and documenting the result.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
中文摘要:MaixCAM2 上使用 camera.Camera 的进程每次退出后,IVPS 在公共池里有时会留下显存块不归还(显式 cam.close() 也一样),驱动不回收已退出进程的块;积累到公共池耗尽后,下一个打开摄像头的进程卡死在 __AX_OSAL_SYNC_wait_interruptible(或 VO init failed),只能整机重启恢复。数据见下。
Environment
- Device: MaixCAM2 (AX630C, 4 GB), factory image, not reflashed
- MaixPy 4.12.5, Axera SDK libs
V3.0.0_20250319114413 - App: a long-running MaixPy app using
camera.Camera(640, 480, image.Format.FMT_RGB888),display.Display(),nn.FaceRecognizer,nn.YOLO11(pose + detect), started from the launcher. A small supervisor process restarts it when it exits.
Symptom
After the app process is restarted a few times (SIGINT, clean exit through app.need_exit()), the next process hangs while initializing the camera / loading models. Its main thread sits in __AX_OSAL_SYNC_wait_interruptible forever; sometimes it fails with RuntimeError: VO init failed instead. Only a full reboot recovers. Killing and restarting the process does not.
What leaks
Reading /proc/ax_proc/pool about 1 s after the old process has exited (no MaixPy process running at that moment) still shows blocks "in use" by IVPS in common pool 1 (BlkSize 2764800, BlkCnt 6), for example:
PoolId IsComm ... BlkSize BlkCnt FreeCnt
1 1 ... 2764800 6 2
Index BlockId VIN IVPS VO ...
0 0x5E001000 0 1 0
2 0x5E001002 0 1 0
3 0x5E001003 0 1 0
5 0x5E001005 0 1 0
These blocks are never returned. Each start also leaves another set of private (non-common) pools behind (the pool id keeps growing), but that alone does not block anything.
Data
Restart loop: SIGINT the app, wait until its heartbeat resumes, repeat. "Residual" = blocks still in use 1 s after the old process exited.
| Round | Residual (original exit path) | Residual (explicit ordered shutdown incl. cam.close() and disp.close()) |
New process |
|---|---|---|---|
| 1 | 2 | 2 | ok |
| 2 | 2 | 2 | ok |
| 3 | 4 | 4 | ok |
| 4 | 5 | 5 | hangs in __AX_OSAL_SYNC_wait_interruptible |
A third run with the same code showed residuals 0, 0, 0, 1, 3 over five rounds, so the leak is intermittent. It looks like it depends on whether frames are still queued in the IVPS output at the moment of exit. It is not related to whether people are in view.
What I checked in MaixCDK
components/vision/port/maixcam2/maix_camera_maixcam2.cpp: Camera::close() calls ax_vi->del_channel() and ax_vi->deinit(), and read() wraps each popped frame in pipeline::Frame(frame, true), so frames handed to the user look released. The leaked blocks are owned by IVPS itself. That suggests frames left in the channel's output queue, or the group itself, are not released on del_channel/deinit, or that the driver does not reclaim blocks of a dead process. The pool/IVPS code is binary-only, so I cannot dig further.
Questions
- Is there an API or a required call sequence (e.g. draining the channel queue, stopping/destroying the IVPS group) that fully returns these blocks before exit?
- Can the driver reclaim pool blocks of a process that has exited?
- Is the common pool configuration (
BlkCnt 6for pool 1) adjustable, to give more headroom?
Workaround in use
After each app exit, a supervisor reads /proc/ax_proc/pool and reboots the device when 3 or more blocks are still held. Any code update is also followed by a full reboot rather than an app restart. This works, but a restart then costs about a minute instead of about 20 seconds.
- Lenguaje dominante
- Python
- Estrellas
- 850
- Forks
- 128
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de sipeed/MaixPy
-
About MaixPy IDE DownloadAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
-
maixhub平台上传应用一个月了还没人审核,什么情况?Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 20/100
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
-
按照官方操作手册转换yolo11模型出现问题Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
-
maixpy 4.12.5 不能支持四分类yolo26Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
Todos los issues de sipeed/MaixPy
Issues similares
-
customer-reported
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Azure/azure-cli#34150 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
community-request
Dificultad 1/5 Menos de una hora Aptitud para principiantes 95/100
NVIDIA-NeMo/Curator#2464 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
weblate-discover crashes with an unhandled FileNotFoundError when the directory does not existAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
WeblateOrg/translation-finder#1099 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
trezor/trezor-firmware#7997 ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día