MLflow tracing silently disabled: setup_mlflow.py sets MLFLOW_CLAUDE_TRACING_ENABLED=false, Stop hook short-circuits
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python
- Domain
- documentation, observability
Research direction
Start with setup_mlflow.py, tests/test_mlflow_tracing.py, and the README tracing section. Review commits b8a06c9 and 66ab612 to understand the intended default before choosing whether tracing should be enabled or opt-in. Align the flag, test names and assertions, documentation, and startup log so the stated behavior matches the deployed behavior.
Written by the indexing model from the issue text.
Description
Symptom
Despite the README promising "every Claude Code session is automatically traced to a Databricks MLflow experiment — zero configuration required", no traces are written. The Stop hook fires, but the upstream MLflow handler returns before processing the transcript.
Root cause
setup_mlflow.py:36 writes the wrong value into ~/.claude/settings.json:
settings["env"]["MLFLOW_CLAUDE_TRACING_ENABLED"] = "false"
This is the env var the upstream mlflow.claude_code package uses as its master enable/disable switch (mlflow-skinny 3.11.1):
mlflow/claude_code/config.py:24—MLFLOW_TRACING_ENABLED = "MLFLOW_CLAUDE_TRACING_ENABLED"mlflow/claude_code/tracing.py:128—is_tracing_enabled()returnsTrueonly when the value is in("true","1","yes")mlflow/claude_code/hooks.py:199—stop_hook_handler()early-returns ifnot is_tracing_enabled():
def stop_hook_handler() -> None:
if not is_tracing_enabled():
response = get_hook_response()
print(json.dumps(response))
return
# ... the transcript-processing path is below this guard
So with "false", the hook prints an empty response and exits — no trace, no error, no log line.
README disagrees with code
The README documents the correct value but the code writes the opposite:
|
MLFLOW_CLAUDE_TRACING_ENABLED|true| Enables Claude Code tracing |
(README.md, "Tracing is configured during app startup" section, line 123.)
Tests lock in the broken behavior
tests/test_mlflow_tracing.py asserts the wrong value, so the test passes against a non-functional setup:
# line 68 and line 143
assert settings["env"]["MLFLOW_CLAUDE_TRACING_ENABLED"] == "false"
A green test suite here is misleading — it certifies the bug.
Repro
- Deploy the app (
make deploy PROFILE=<profile>). - Open the app, start a Claude Code session, run a few prompts.
- Type
exitto fire the Stop hook. - Open the MLflow experiment at
/Users/{app_owner}/coding-agents.
Expected: A trace per session, with prompts/tool calls/outputs visible.
Actual: Experiment is empty (or contains only traces from before this regression landed).
Suggested fix
- settings["env"]["MLFLOW_CLAUDE_TRACING_ENABLED"] = "false"
+ settings["env"]["MLFLOW_CLAUDE_TRACING_ENABLED"] = "true"
Plus update the two test assertions in tests/test_mlflow_tracing.py (lines 68 and 143).
Notes
- The OTEL endpoint override on the next line (
OTEL_EXPORTER_OTLP_ENDPOINT = "") is unrelated and looks correct — it disables the container's OTLP collector so MLflow uses its native exporter. git blame setup_mlflow.pywould show when"false"was introduced and whether it predates the upstream gating change in mlflow-skinny 3.11.x — worth a glance to decide whether to backport the fix to older releases.
Update — root cause was intentional
git log on setup_mlflow.py shows the disable was intentional. Commit b8a06c9 (2026-03-28) is titled "fix: always create fresh session on page load, disable MLflow tracing by default" and explicitly flips the flag from "true" to "false".
So this is deliberate-disable + stale-docs, not a stray flag. Two ways to resolve:
- (A) Re-enable — flip back to
"true", verify whatever motivatedb8a06c9is gone (the commit message bundled "fresh session on page load" with the disable, so the underlying issue may have been addressed alongside). - (B) Keep disabled, update docs — leave
"false", rewrite the README "Tracing is configured during app startup" section to say tracing is opt-in, restore the env-driven override that existed in earlier history (commit66ab612), and renametest_tracing_enabledto match.
Either way the README, the test name, and the flag value need to align with whichever direction is chosen.
Empirical confirmation (2026-05-06, v0.18.1 on daveok)
Deployed databrickslabs/main (v0.18.1, commit dc14bf8) to a fresh Databricks Apps workspace, completed setup, and verified:
1. The flag is "false" in the running container. Read directly from the deployed ~/.claude/settings.json:
FLAG= false EXP= /Users/<owner>/coding-agents OTEL= ''
2. The MLflow experiment doesn't exist. Querying GET /api/2.0/mlflow/experiments/get-by-name?experiment_name=/Users/<owner>/coding-agents returns:
{"error_code": "RESOURCE_DOES_NOT_EXIST",
"message": "Node /Users/<owner>/coding-agents does not exist."}
So nothing is being written. MLflow auto-creates experiments on first write — since the Stop hook short-circuits before any write, no experiment ever materialises. The user has no breadcrumb to follow when they go looking for the traces the README promised them.
Adjacent finding: misleading log line
setup_mlflow.py:62 prints "MLflow tracing enabled: experiment={...}" even though it just wrote MLFLOW_CLAUDE_TRACING_ENABLED="false" two lines earlier. The startup log says "enabled" — the user has every reason to believe tracing is working. This makes the silent-failure mode worse, not better.
Whichever direction is chosen for the master fix, this print message should match reality (or be removed).
- Dominant language
- Python
- Stars
- 40
- Forks
- 11
- Avg merge
- 1m
- Merged PRs (30d)
- 1
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from databrickslabs/coding-agents-databricks-apps
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
-
Difficulty 3/5 1-2 days Newbie friendliness 68/100
-
Difficulty 4/5 3-5 days Newbie friendliness 55/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
All issues in databrickslabs/coding-agents-databricks-apps
Similar issues
-
agent-ready documentation needs-triage
Difficulty 1/5 1-3 hours Newbie friendliness 88/100
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 91/100
-
workflow-status page template still says reusable workflows are "triggered only by workflow_call:" Open
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
instance instance add
Difficulty 1/5 Under an hour Newbie friendliness 72/100
searxng/searx-instances#939 · 1 comment ·
-
area-deployment area-integrations triage:bot-seen
Difficulty 2/5 Half a day Newbie friendliness 86/100