Feature Request: Support for Financial RAG Dataset Eval Pipelines with Domain-Specific Scorers
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 25/100
- issue の種類
- 機能追加
- 明瞭さ
- 説明が足りない
- 活発さ
- 静か
- 技術スタック
- python
調査の方向性
まず、issue で説明されている EvalCase スキーマと scorer 登録 API を確認し、それらを既存の Braintrust Eval の使用方法と比較します。実装前に、context のサポート、数値同値性のスコアリング、デコレータベースの登録のどれに絞るかを決めてスコープを狭める必要があります。done は、選択した提案とその検証カバレッジによって定義する必要があります。
索引モデルが issue の本文から書いたものです。
説明
Summary
I've been building a financial RAG evaluation framework (FinRAG-Eval) and ran into friction when trying to use braintrust-sdk-python for domain-specific eval pipelines over financial documents (10-Ks, earnings transcripts, SEC filings).
Problem
When running evals on financial RAG outputs, the default scorer setup doesn't map well to domain-specific correctness signals. Specifically:
- No built-in support for numerical/unit-aware comparison — financial answers often contain figures like
$3.2Bvs3.2 billion. Current string-match scorers treat these as mismatches. - No dataset schema for context-grounded financial Q&A — when loading eval datasets from Braintrust's dataset store, there's no documented convention for attaching retrieved context chunks (needed for faithfulness scoring).
- Custom scorer registration is verbose — adding a domain scorer (e.g., a ROUGE-F1 scorer or a financial entity extractor) requires wrapping functions manually with no type hints or schema validation.
Proposed Solution
- Add a
contextfield to the standardEvalCaseschema (alongsideinput,expected,metadata) so retrieved chunks can be passed through to scorers natively. - Provide a
NumericEquivalenceScorerthat normalizes units (B/M/K, $, %) before comparing. - Allow scorer registration via a decorator pattern (
@braintrust.scorer) similar to how pytest fixtures work — this would reduce boilerplate significantly.
Example Use Case
@braintrust.scorer
def financial_faithfulness(output: str, context: list[str]) -> Score:
# Check if numerical claims in output are grounded in context
...
return Score(name="financial_faithfulness", score=0.87)
await Eval(
"FinRAG-Eval",
data=financial_qa_dataset,
task=rag_pipeline,
scores=[financial_faithfulness, NumericEquivalenceScorer()],
)
Context
I'm building this as part of finrag-eval, a framework for evaluating LLM outputs over financial documents. Happy to contribute a PR for the scorer decorator pattern or the context field addition if the team is open to it.
References
- Braintrust Eval Docs
- Related: DeepEval's
LLMTestCasehas aretrieval_contextfield that works similarly
- 主要言語
- Python
- スター
- 20
- フォーク
- 18
- 平均マージ
- 21時間 8分
- マージ済み PR(30日)
- 81
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
braintrustdata/braintrust-sdk-python のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
braintrustdata/braintrust-sdk-python#797 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
braintrustdata/braintrust-sdk-python#774 ·
メンテナーはふだん 1 日以内に返信
-
Migrate legacy HTTPConnection callers to the policy-aware Transport incrementally対応中かも @AbhiPrasad が 1 日前に担当しました。 オープンpython
braintrustdata/braintrust-sdk-python#839 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
new-integration
難易度 4/5 3〜5日 初心者へのやさしさ 68/100
braintrustdata/braintrust-sdk-python#808 ·
メンテナーはふだん 1 日以内に返信
-
new-integration
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
braintrustdata/braintrust-sdk-python#807 ·
メンテナーはふだん 1 日以内に返信
braintrustdata/braintrust-sdk-python の issue をすべて見る
似ている issue
-
#bug
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
apache/superset#44923 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
lawndoc/stack-back#123 ·
-
Add: Entuneオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
AbdelStark/awesome-typesafe-jev#187 ·
メンテナーはふだん 1 日以内に返信
-
bug good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
repowise-dev/repowise#2966 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 2 日以内に返信