Implement the unified attention interpretation API for similar models
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- huggingface, python, pytorch, scikit-learn
- Domain
- ai, machine-learning
Research direction
Start by reading the NeuroX codebase and the linked HuggingFace and Captum material, then compare the BERT and RoBERTa interpretation paths. Define the model scope and API boundary before implementation; done should mean a unified interpretability interface for the selected crucial models, with encoder and encoder-decoder coverage only if explicitly included.
Written by the indexing model from the issue text.
Description
duration: scalable, can be both 170 and 340 hours
difficulty: challenging
mentor: @oserikov , TBD
requirements:
- pytorch
- sklearn
- python engineering code, OOP, etc.
- experience with Transformer Language models
useful links:
- NeuroX codebase
- Bert re-invents the classical NLP pipeline
- Captum
Idea Description:
While HuggingFace quickly became the standard way to publish language models, several architectural trade-offs have been made to support the quick growth of the models' zoo. This resulted in several theoretically similar models being implemented by different teams, thus e.g. several alternative implementations of self-attentive transformers arose. While refactoring the whole zoo of models seems to be far from the accessible task, the interpretability community is forced to provide unification wrappers for handling such dissimilarities in similar models. The task is to provide a reasonable trade-off with the refactoring of the crucial models and providing the unified wrappers, and thus bring the unified interpretability API to the crucial HuggingFace models.
We could see this task from two prospects. First, one could unify the interpretability API of the sibling models such as BERT and RoBERTa . Second, one could think about bringing the unified interface to interpret and compare encoder models with e.g. encoder-decoder ones, allowing to study the similarities and distinctiveness in their behavior.
Coding Challenge
WIP
- Dominant language
- No language data
- Stars
- 1
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from bigscience-workshop/interpretability-ideas
-
Difficulty 5/5 Over a week Newbie friendliness 20/100
-
Difficulty 5/5 Over a week Newbie friendliness 10/100
-
Difficulty 5/5 Over a week Newbie friendliness 20/100
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
-
Difficulty 5/5 Over a week Newbie friendliness 10/100
All issues in bigscience-workshop/interpretability-ideas
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
NousResearch/hermes-plugin-claude-subscription-directsdk#67 ·
Maintainers usually reply within 1 day
-
area:dictation bug P2
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
uttrflow/uttrflow-swift#2539 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
comp/tools P3 type/bug
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
NousResearch/hermes-agent#126059 ·
Maintainers usually reply within 1 day
-
factory-active factory-automatic harness/codex task-bug-reproduction-success task-identify-harness-labels-done task-identify-issue-type-done
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
vercel/ai#21582 · 3 comments ·
Maintainers usually reply within 1 day