Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Partial formula execution

Open
#65 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Feature
Clarity
Needs clarification
Activity status
Stale
Tech stack
python

Research direction

Start by reading issue #64 and its prototype, then trace the variable execution path around the proposed eligible and formula entry points. Compare the current behavior with the Massachusetts-only example and identify how to avoid breaking changes. Done means a considered implementation that evaluates formulas only for eligible entities without the bugs noted in #64.

Written by the indexing model from the issue text.

Description

enhancement

Although this might be relevant to Core, I suspect it'd need a much longer discussion to avoid breaking changes, so filing here with a view to implementing as a patch. There have been a few attempts already in #64 , but with some bugs so I though I'd sketch out the cleanest implementation here.

The problem

Some variables are only relevant to a small subset of the population. For example, Massachusetts income tax only needs to be calculated for people and groups in Massachusetts, and not the rest of the population. Right now, we implement the tax as simply zero for those other people, but this causes wasted computation time and space for 98% of entities, because NumPy vectorised operations happen regardless of the retrospective filter at the end.

The solution

We could have the following variable definition:

class ma_tax(Variable):
  value_type = float
  label = "MA income tax"
  definition_period = YEAR
  unit = TaxUnit
  
  def eligible(tax_unit, period, parameters):
    return tax_unit.household("state_code", period) == "MA"
  
  def formula(tax_unit, period, parameters):
    ...

eligible is run first to determine the relevant subset of the population, and then the main formula next. This will be much more efficient iff the formula is much more complex than eligible.

#64 has a prototype of the implementation, but it's buggy and needs more thought. I think there's a clean way to do this, intercepting the population passed to the formula to only return the subset values.

cc @MattiSG, @MaxGhenis, @rickecon

Dominant language
Python
Stars
1
Forks
1
PR merge metrics
No merged PRs in 30d

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from PolicyEngine/openfisca-tools

All issues in PolicyEngine/openfisca-tools

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.