Partial formula execution
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- python
- Domain
- backend, performance
Research direction
Start by reading issue #64 and its prototype, then trace the variable execution path around the proposed eligible and formula entry points. Compare the current behavior with the Massachusetts-only example and identify how to avoid breaking changes. Done means a considered implementation that evaluates formulas only for eligible entities without the bugs noted in #64.
Written by the indexing model from the issue text.
Description
Although this might be relevant to Core, I suspect it'd need a much longer discussion to avoid breaking changes, so filing here with a view to implementing as a patch. There have been a few attempts already in #64 , but with some bugs so I though I'd sketch out the cleanest implementation here.
The problem
Some variables are only relevant to a small subset of the population. For example, Massachusetts income tax only needs to be calculated for people and groups in Massachusetts, and not the rest of the population. Right now, we implement the tax as simply zero for those other people, but this causes wasted computation time and space for 98% of entities, because NumPy vectorised operations happen regardless of the retrospective filter at the end.
The solution
We could have the following variable definition:
class ma_tax(Variable):
value_type = float
label = "MA income tax"
definition_period = YEAR
unit = TaxUnit
def eligible(tax_unit, period, parameters):
return tax_unit.household("state_code", period) == "MA"
def formula(tax_unit, period, parameters):
...
eligible is run first to determine the relevant subset of the population, and then the main formula next. This will be much more efficient iff the formula is much more complex than eligible.
#64 has a prototype of the implementation, but it's buggy and needs more thought. I think there's a clean way to do this, intercepting the population passed to the formula to only return the subset values.
cc @MattiSG, @MaxGhenis, @rickecon
- Dominant language
- Python
- Stars
- 1
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from PolicyEngine/openfisca-tools
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
All issues in PolicyEngine/openfisca-tools
Similar issues
-
namespace operations
Difficulty 1/5 Under an hour Newbie friendliness 72/100
EclipseFdn/open-vsx.org#14043 ·
Maintainers usually reply within 1 day
-
netbox status: needs triage type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
netbox-community/netbox#23376 ·
Maintainers usually reply within 1 day
-
feedback simulation workshop
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
githubnext/gh-aw-workshop#4455 ·
Maintainers usually reply within 1 day
-
Triage 🩺
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOpenneeds-triage
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
krkn-chaos/krkn#1627 · 1 comment ·
Maintainers usually reply within 1 day