Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[API] `flashdreams.core` Refactor

Open
#478 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Refactor
Clarity
Needs clarification
Activity status
Quiet
Tech stack
python

Research direction

Start by reviewing dependent issues #473 and #482, then inspect flashdreams/flashdreams/core/attention, flashdreams/core/distributed, flashdreams/infra/acceleration, and the proposed flashdreams.accelerated work. Map the current modules, acceleration, distributed features, and StreamInferencePipeline before changing layout. Done means the layered core-to-pipeline-to-runtime structure is in place and the obsolete flashdreams/infra code and pipeline are removed.

Written by the indexing model from the issue text.

Description

Priority-P1 Type-epic

Issue Description

Current flashdreams spreads the "core level" feature in multiple places under flashdreams/flashdreams. We need to refactor/reorganize the code layout to have a layered structure (flashdreams.core -> flashdreams.pipeline -> flashdreams.runtime)

Proposed Solution

The following features should go into the new flashdreams.core:

  • flashdreams.core.modules: accelerated flashdreams modules with their Triton kernel. Currently, they are in https://github.com/NVIDIA/flashdreams/tree/main/flashdreams/flashdreams/core/attention (special rope module & rope triton kernel) and my new flashdreams.accelerated PR (WIP)
  • flashdreams.core.acceleration: CUDA graph, prewarm, and future flashdreams auto tune system. These feature are currently in flashdreams/infra/acceleration
  • flashdreams.core.distributed: Distributed related features. Including context parallel (currently in flashdreams/core/attention/cp.py and flashdreams/core/distributed), rank orchestration (currently in flashdreams/core/distributed)

We also need to remove all the old StreamInferencePipeline in favor of the composable inference pipeline once implemented. This should mean that we no longer need flashdreams/infra after the refactor.

Dependency

This issue is gated by #473 and #482

Dominant language
Python
Stars
510
Forks
61
Avg merge
3d 4h
Merged PRs (30d)
49

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/flashdreams

All issues in NVIDIA/flashdreams

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.