chore!: Automate data model generation from upcoming CycloneDX 2.0 modularized specification
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Refactoring
- Chiarezza
- Da chiarire
- Stato di attività
- Tranquilla
- Stack tecnologico
- python
- Ambito
- build-system, tooling
Direzione di ricerca
Iniziare con lo schema CycloneDX 2.0 in schema/2.0/model e revisionare gli strumenti di generazione indicati e il lavoro di proof-of-concept. Definire come la pre-elaborazione, la generazione del codice, la post-elaborazione e l’output committato si integrano nel processo di build o release del progetto. Done include modelli deterministici, documentazione ed esempi aggiornati, test dello schema e di enum, e una guida alla transizione o limitazioni della migrazione documentate.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Description
Currently, the data models in this library are largely written and maintained manually. While this approach has worked so far, it is time-consuming and requires significant effort for both implementation and review. This effort could be better invested in feature development and bug fixing. (Also the issue is that the models code also implement the business logic, like property constraints, etc.)
With the upcoming CycloneDX 2.0 specification, a modularized and machine-readable format will be introduced. This presents an opportunity to rethink how data models are created and maintained in this project.
Reference (work in progress):
- PR https://github.com/CycloneDX/cyclonedx-python-lib/issues
- modularized schema https://github.com/CycloneDX/specification/tree/2.0-dev/schema/2.0/model
Problem
- Data models are mostly handwritten and mix in the business logic
- High maintenance overhead
- Repetitive work for contributors
- Slows down development velocity due to review effort
Proposal
Leverage the machine-readable specification planned for CycloneDX 2.0 to introduce static code generation for data models (not business logic).
This would involve:
- Parsing the official CycloneDX specification (once available in machine-readable form)
- Generating Python data models automatically
- Integrating generation into the build or release process
- Minimizing manual intervention for future spec updates
There have already been proof-of-concept implementations demonstrating that automated generation of data models from the specification is feasible. These approaches should be revisited, consolidated, and applied as part of this effort.
Pipeline:
flowchart TB
A["Input: CycloneDX JSON Schema files"]
B["Preprocessing (if needed)"]
C["Code Generation"]
D["Post-processing: formatting, code style, adjustments"]
subgraph E["Result: Generated Python code"]
direction LR
E1["Data Models"]
E2["Factories"]
E3["Serialization-specific Normalizers"]
end
A --> B --> C --> D --> E
Philosophy: decouple logic fro mdata
- Models are pure dumb typed data classes.
- Factories implement business logic for instancing data classes with the propeties' constrinats defined in the schema
- normalizers turn data classes into objects that can be handed over to serializers
- serializers take prepared objects and render XML/JSON/PB
possible file layput in the distributable:
Core
V1
v2
Models
Cyclonedx
Modules
Common
Normalizers
Serializes
Possible Tools / Libraries
The following tools could be evaluated as part of this effort:
-
datamodel-code-generator — MIT
https://pypi.org/project/datamodel-code-generator/ -
pydantic — MIT
https://pypi.org/project/pydantic/ -
dataclasses-json — MIT
https://pypi.org/project/dataclasses-json/ -
dacite — MIT
https://pypi.org/project/dacite/ -
marshmallow — MIT
https://pypi.org/project/marshmallow/ -
marshmallow-jsonschema — MIT
https://pypi.org/project/marshmallow-jsonschema/ -
jsonschema (validation, not models) — MIT
https://pypi.org/project/jsonschema/ -
quicktype — Apache 2.0
https://pypi.org/project/quicktype/ -
genson (schema generator, reverse direction) — MIT
https://pypi.org/project/genson/
Community Input
Community discussions have already suggested evaluating tools such as:
- datamodel-code-generator
- json-schema-to-pydantic (https://pypi.org/project/json-schema-to-pydantic/)
- jambo (https://pypi.org/project/jambo/)
and
- de/serialization with cattrs (https://pypi.org/project/cattrs/)
These should be considered as primary candidates during evaluation.
see discussions:
Expected Benefits
- Significant reduction in maintenance effort
- Improved consistency across models
- Faster adoption of new specification versions
- More time available for feature development and bug fixing
Considerations / Open Questions
- What format will the machine-readable spec be published in (e.g., JSON Schema, OpenAPI, etc.)?
- JSON Schema it is
- Should generated code be committed or generated at build time?
- decision: generated before build time, and commited to the repo
- How to handle custom logic or extensions on top of generated models?
- Backward compatibility with CycloneDX 1.x
- easy path: breaking change in the library, and only support 2.0 from then on
Additional Context
This proposal aligns with the direction of CycloneDX 2.0, which aims to make the specification more modular and tooling-friendly. Taking advantage of this early could significantly improve long-term maintainability of this library.
Note: This issue is intended as a meta-ticket to collect related sub-tasks and track overall implementation efforts.
Checklist
- have the models created in a deterministic way from schema
- generate the docs
- have examples updated
- have tests
- test with all schema test cases
- tests all enum completeness
- write a transition guide in the docs
- section "Upgrading to vXXX"
- we might not be able to provide a proper guide at all, since basically everything is rewrite.
- we might point to the github discussions - section FAQ - to pride at least some help
- Lingua principale
- Python
- Stelle
- 117
- Fork
- 67
- Merge medio
- 21h 9m
- PR unite (30g)
- 3
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di CycloneDX/cyclonedx-python-lib
-
[PERF] Quadratic (O(N^2)) serialization time for large BOMs — `Bom.validate()` → `register_dependency()` linear scanForse già presa @inspired-geek l’ha presa 109 giorni fa. Apertaperformance
Difficoltà 3/5 1-2 giorni Idoneità per principianti 36/100
CycloneDX/cyclonedx-python-lib#1006 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
tests: test all model enumsForse di nuovo libera @jkowalleck l’ha presa 125 giorni fa e non c’è nessuna pull request aperta. ApertaQA
CycloneDX/cyclonedx-python-lib#991 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
feat(deps)!: make all de/serialization libraries optionalForse di nuovo libera @Simoh23999 l’ha presa 73 giorni fa e non c’è nessuna pull request aperta. Apertabreaking change dependencies
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
CycloneDX/cyclonedx-python-lib#979 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
feat: Add support for Component signatureForse già presa @wiebe-vandendriessche l’ha presa 124 giorni fa. Apertaenhancement help wanted schema 1.4
Difficoltà 3/5 1-2 giorni Idoneità per principianti 58/100
CycloneDX/cyclonedx-python-lib#978 · 4 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
chore: have coverage uploaded consitentlyForse di nuovo libera @jkowalleck l’ha presa 171 giorni fa e non c’è nessuna pull request aperta. Apertachore
CycloneDX/cyclonedx-python-lib#966 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di CycloneDX/cyclonedx-python-lib
Issue simili
-
Claiming namespace `jft63`Apertanamespace operations
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
EclipseFdn/open-vsx.org#14043 ·
I maintainer di solito rispondono entro 1 giorno
-
netbox status: needs triage type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
netbox-community/netbox#23376 ·
I maintainer di solito rispondono entro 1 giorno
-
feedback simulation workshop
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
githubnext/gh-aw-workshop#4455 ·
I maintainer di solito rispondono entro 1 giorno
-
Triage 🩺
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitApertaneeds-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 77/100
krkn-chaos/krkn#1627 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno