Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Feature Request: Support for Financial RAG Dataset Eval Pipelines with Domain-Specific Scorers

オープン
#492 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
25/100
issue の種類
機能追加
明瞭さ
説明が足りない
活発さ
静か
技術スタック
python

調査の方向性

まず、issue で説明されている EvalCase スキーマと scorer 登録 API を確認し、それらを既存の Braintrust Eval の使用方法と比較します。実装前に、context のサポート、数値同値性のスコアリング、デコレータベースの登録のどれに絞るかを決めてスコープを狭める必要があります。done は、選択した提案とその検証カバレッジによって定義する必要があります。

索引モデルが issue の本文から書いたものです。

説明

Summary

I've been building a financial RAG evaluation framework (FinRAG-Eval) and ran into friction when trying to use braintrust-sdk-python for domain-specific eval pipelines over financial documents (10-Ks, earnings transcripts, SEC filings).

Problem

When running evals on financial RAG outputs, the default scorer setup doesn't map well to domain-specific correctness signals. Specifically:

  1. No built-in support for numerical/unit-aware comparison — financial answers often contain figures like $3.2B vs 3.2 billion. Current string-match scorers treat these as mismatches.
  2. No dataset schema for context-grounded financial Q&A — when loading eval datasets from Braintrust's dataset store, there's no documented convention for attaching retrieved context chunks (needed for faithfulness scoring).
  3. Custom scorer registration is verbose — adding a domain scorer (e.g., a ROUGE-F1 scorer or a financial entity extractor) requires wrapping functions manually with no type hints or schema validation.

Proposed Solution

  • Add a context field to the standard EvalCase schema (alongside input, expected, metadata) so retrieved chunks can be passed through to scorers natively.
  • Provide a NumericEquivalenceScorer that normalizes units (B/M/K, $, %) before comparing.
  • Allow scorer registration via a decorator pattern (@braintrust.scorer) similar to how pytest fixtures work — this would reduce boilerplate significantly.

Example Use Case

@braintrust.scorer
def financial_faithfulness(output: str, context: list[str]) -> Score:
    # Check if numerical claims in output are grounded in context
    ...
    return Score(name="financial_faithfulness", score=0.87)

await Eval(
    "FinRAG-Eval",
    data=financial_qa_dataset,
    task=rag_pipeline,
    scores=[financial_faithfulness, NumericEquivalenceScorer()],
)

Context

I'm building this as part of finrag-eval, a framework for evaluating LLM outputs over financial documents. Happy to contribute a PR for the scorer decorator pattern or the context field addition if the team is open to it.

References

  • Braintrust Eval Docs
  • Related: DeepEval's LLMTestCase has a retrieval_context field that works similarly
主要言語
Python
スター
20
フォーク
18
平均マージ
21時間 8分
マージ済み PR(30日)
81

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

braintrustdata/braintrust-sdk-python のほかの issue

braintrustdata/braintrust-sdk-python の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。