[Bug] Internal execution failure due to container image error after Helm chart deployment
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 38/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Tranquilo
- Stack tecnológico
- docker, helm, kubernetes
- Área
- build-system, devops, infrastructure
Línea de trabajo
Start by reproducing the failure with ghcr.io/microsoft/jrtc-apps/grc:latest, running volk_profile -v and the GRC flowgraph in the Helm deployment. Compare the image build configuration and CPU targets with a locally built image on the Intel Xeon Gold 6148, and check the provided gNB image similarly. Done means the container image avoids the segmentation fault and both GRC and gNB start successfully on the reported host.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Hi all,
First of all, since I have already successfully built and executed everything on a bare-metal environment, I am confident that there are no issues with the source code itself. However, I am concerned that the provided built image (especially on grc, and srsgnb) might have been compiled too strictly (e.g., heavily optimized for specific newer CPU architectures).
1. Environment
Image: ghcr.io/microsoft/jrtc-apps/grc:latest
GNU Radio Version: 3.10.1.1
Host OS: Ubuntu 22.04.5 LTS (Kernel 5.15.0-119-generic)
CPU: Intel(R) Xeon(R) Gold 6148 (Supports AVX-512)
2. The Core Issue
The GRC process hangs indefinitely after printing the following log:
[INFO] Starting flowgraph....
When running volk_profile -v inside the container, it immediately crashes:
command terminated with exit code 139 (Segmentation Fault)
**Our Suspicion: ** This indicates a failure within the VOLK library, likely during SIMD kernel selection/dispatching. We suspect a potential -march=native optimization issue during the image build process. If the image was compiled on a different CPU architecture (e.g., a newer generation Xeon), the specific instructions might be incompatible with our 1st Gen Xeon Gold processor, leading to this crash.
3. Related Binary Issues (gNB Case)
We encountered a similar compatibility issue with the gNB image provided in the Helm chart, where the original binary crashed upon startup.
We were only able to resolve the gNB issue by replacing the container's binary with one built locally on the same host.
This consistent failure across both gNB and GRC images strongly suggests a general binary incompatibility with our environment.
4. Verified Configurations & Troubleshooting
We have ruled out permission and networking issues by verifying the following:
-
Applied yaml file ( privileged: true, hostIPC: true, and seccomp: Unconfined to the pod. )
-
Verified ZMQ connectivity (E2E).
-
Subscriber successfully registered in Open5GS DB.
5. Our Situation and Concerns
We have transitioned to a Kubernetes-based setup following your recommendation to achieve stability. However, encountering these fundamental binary issues at the initial stage is quite unexpected.
While we could technically proceed by manually building and replacing all binaries locally, we are deeply concerned that such a workaround will lead to future compatibility risks and massive maintenance overhead. To ensure a stable and sustainable setup as intended by your project, we believe a fundamental fix at the image level is necessary.
6. Requests
-
Could you verify if these images were built with specific CPU optimizations (e.g., AVX-512 for newer architectures) that might be unstable on certain Xeon architectures?
-
Could you provide a more generic image (e.g., compiled with AVX2 or Generic targets) or advise on a definitive way to resolve these instruction set mismatches?
We really want to avoid relying on local manual builds for long-term maintenance. Looking forward to your thoughts on this.
- Lenguaje dominante
- Python
- Estrellas
- 14
- Forks
- 13
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/jrtc-apps
-
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
-
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
-
Add basic tests to the CIQuizá libre de nuevo @doctorlai-msrc la tomó hace 441 días y no hay ningún pull request abierto. Abierto
Todos los issues de microsoft/jrtc-apps
Issues similares
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
letsencrypt/cp-cps#353 ·
-
Marble Madness II is missingAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
DOI-USGS/pywatershed#421 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
python-pillow/Pillow#10087 · 1 comentario ·
Los mantenedores suelen responder en 1 día