Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

chore!: Automate data model generation from upcoming CycloneDX 2.0 modularized specification

Aperta
#955 3 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Refactoring
Chiarezza
Da chiarire
Stato di attività
Tranquilla
Stack tecnologico
python

Direzione di ricerca

Iniziare con lo schema CycloneDX 2.0 in schema/2.0/model e revisionare gli strumenti di generazione indicati e il lavoro di proof-of-concept. Definire come la pre-elaborazione, la generazione del codice, la post-elaborazione e l’output committato si integrano nel processo di build o release del progetto. Done include modelli deterministici, documentazione ed esempi aggiornati, test dello schema e di enum, e una guida alla transizione o limitazioni della migrazione documentate.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

breaking change chore enhancement schema 2.0
Description

Currently, the data models in this library are largely written and maintained manually. While this approach has worked so far, it is time-consuming and requires significant effort for both implementation and review. This effort could be better invested in feature development and bug fixing. (Also the issue is that the models code also implement the business logic, like property constraints, etc.)

With the upcoming CycloneDX 2.0 specification, a modularized and machine-readable format will be introduced. This presents an opportunity to rethink how data models are created and maintained in this project.

Reference (work in progress):


Problem
  • Data models are mostly handwritten and mix in the business logic
  • High maintenance overhead
  • Repetitive work for contributors
  • Slows down development velocity due to review effort

Proposal

Leverage the machine-readable specification planned for CycloneDX 2.0 to introduce static code generation for data models (not business logic).

This would involve:

  • Parsing the official CycloneDX specification (once available in machine-readable form)
  • Generating Python data models automatically
  • Integrating generation into the build or release process
  • Minimizing manual intervention for future spec updates

There have already been proof-of-concept implementations demonstrating that automated generation of data models from the specification is feasible. These approaches should be revisited, consolidated, and applied as part of this effort.

Pipeline:

flowchart TB
    A["Input: CycloneDX JSON Schema files"]
    B["Preprocessing (if needed)"]
    C["Code Generation"]
    D["Post-processing: formatting, code style, adjustments"]

    subgraph E["Result: Generated Python code"]
        direction LR
        E1["Data Models"]
        E2["Factories"]
        E3["Serialization-specific Normalizers"]
    end

    A --> B --> C --> D --> E

Philosophy: decouple logic fro mdata

  • Models are pure dumb typed data classes.
  • Factories implement business logic for instancing data classes with the propeties' constrinats defined in the schema
  • normalizers turn data classes into objects that can be handed over to serializers
  • serializers take prepared objects and render XML/JSON/PB

possible file layput in the distributable:

Core
  V1
  v2
    Models
      Cyclonedx
      Modules
        Common
   Normalizers
   Serializes

Possible Tools / Libraries

The following tools could be evaluated as part of this effort:


Community Input

Community discussions have already suggested evaluating tools such as:

and

These should be considered as primary candidates during evaluation.

see discussions:


Expected Benefits
  • Significant reduction in maintenance effort
  • Improved consistency across models
  • Faster adoption of new specification versions
  • More time available for feature development and bug fixing

Considerations / Open Questions
  • What format will the machine-readable spec be published in (e.g., JSON Schema, OpenAPI, etc.)?
    • JSON Schema it is
  • Should generated code be committed or generated at build time?
    • decision: generated before build time, and commited to the repo
  • How to handle custom logic or extensions on top of generated models?
  • Backward compatibility with CycloneDX 1.x
    • easy path: breaking change in the library, and only support 2.0 from then on
Additional Context

This proposal aligns with the direction of CycloneDX 2.0, which aims to make the specification more modular and tooling-friendly. Taking advantage of this early could significantly improve long-term maintainability of this library.


Note: This issue is intended as a meta-ticket to collect related sub-tasks and track overall implementation efforts.


Checklist

  • have the models created in a deterministic way from schema
  • generate the docs
  • have examples updated
  • have tests
    • test with all schema test cases
    • tests all enum completeness
  • write a transition guide in the docs
    • section "Upgrading to vXXX"
    • we might not be able to provide a proper guide at all, since basically everything is rewrite.
    • we might point to the github discussions - section FAQ - to pride at least some help
Lingua principale
Python
Stelle
117
Fork
67
Merge medio
21h 9m
PR unite (30g)
3

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di CycloneDX/cyclonedx-python-lib

Tutte le issue di CycloneDX/cyclonedx-python-lib

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.