Chatbot Engine: Dynamic, configurable conversations
维护者通常 2 天内回复
@Ayush8923 已经在做这个了。
开始于 2026年10月5日。
评估
这个 Issue 还没有评估数据。
描述
Describe the current behavior
As a first step (#1210), we built a read-only /agent endpoint: a single, static LangGraph agent that answers natural-language questions about a tenant own eval runs, datasets, etc. It is "static" in two senses:
- The graph itself is fixed in code, one
agent_node/tool_executor_nodeloop, hand-written, the same for every caller. - The toolset is a fixed, hand-written registry of read-only GET wrappers (
list_evaluation_runs,get_collection, etc.), the agent picks which of these to call and in what order, but the set of available tools and the conversation shape never change per request.
That's the right shape for "ask a question, get an answer" over existing data. It is not the right shape for the next use case below, where the conversation itself needs to be admin-configurable, stateful across many turns/days, and able to take write-side actions (set a reminder, etc.), not just read.
Describe the enhancement you'd like
We want to generalize from "one static read-only agent" to a dynamic, admin-configurable conversational bot flow engine, the foundation for use cases like NGO onboarding on Glific:
When an NGO onboards, today they fill out a static form. We want to replace this with a chatbot that asks a configurable sequence of questions one at a time, analyzes each answer, and, if the user doesn't answer a question, retries up to an admin-configured count before applying an admin-configured fallback (skip the question, end the conversation, escalate, etc.). The flow (questions, retry/skip rules, branching) must be fully dynamic/admin-authored, not hand-coded per bot. The bot also has access to an admin-selectable catalogue of tools (Recurring Reminder, One-off Reminder, Delete Reminder, etc.) that it can call autonomously, the same way #1210's agent calls read-only Kaapi endpoints.
Reference: Open Chat Studio (OCS)
We looked at OCS (github.com/dimagi/open-chat-studio) again as a reference, not a blueprint:
- OCS "Pipelines" are a visual node graph (nodes + edges) that compiles directly into a LangGraph StateGraph at runtime, i.e. the graph is built from a stored config, not hand-written per bot. We want to adopt this shape.
- Node types are typed Pydantic classes whose fields double as the admin-facing config form (prompt text, which tools are enabled for this node, max results, etc.) — a clean pattern for "admin picks from a fixed catalogue" without a generic/untyped config blob.
- Retry/skip-after-N-no-answer is not a graph loop construct in OCS, it lives in per-session state (a counter) plus a conditional router, decoupled from the tool/LLM machinery. We want the same separation.
- Reminders/scheduling in OCS are a separate Celery-driven trigger system (
TimeoutTrigger,StaticTrigger,ScheduledTrigger), not something that blocks inside the live conversation graph. A "set a reminder" tool call should enqueue a scheduled action and return immediately, not hold the graph open. - OCS persists conversation state via its own DB models, not LangGraph checkpointer. We think we should diverge here (see below), our eval-iteration loop already proves out langgraph-checkpoint-postgres for exactly this kind of long-lived, resumable, multi-day conversation, so we'd rather reuse that than build a parallel persistence layer.
- OCS versions every
pipeline/trigger/custom-action(working vs. published), so an admin can edit a flow without affecting live conversations. We want this too, and it maps directly onto Kaapi's existingConfig / ConfigVersionpattern (already used by/llm/call) rather than needing a new concept.
Architecture
Flow definition, dynamic graph, built from a versioned config
- New
FlowConfig/FlowConfigVersiontables, following the existing Config / ConfigVersion pattern. The version's JSON blob declares nodes + edges:
{
"nodes": [
{"id": "q1", "type": "question", "prompt": "What is your NGO's name?",
"max_retries": 2, "on_exhausted": "skip"},
{"id": "q2", "type": "question", "prompt": "Registration number?",
"max_retries": 2, "on_exhausted": "end_conversation"},
{"id": "tools_step", "type": "llm_with_tools",
"tools": ["recurring_reminder", "one_off_reminder"]}
],
"edges": [{"from": "q1", "to": "q2"}, {"from": "q2", "to": "tools_step"}]
}
- A graph builder maps each node's "type" to a typed Python node class and compiles a LangGraph StateGraph from the declared nodes/edges at runtime. Adding a new node type is a code change (one Python class); authoring a new flow is a config change only.
- The LangGraph
checkpointeronly exists to resume a paused conversation; it is not a queryable business record. A separateFlowSessiontable
flow_config_id: UUID # which flow
flow_config_version: int # version
contact_ref: str # Glific/WhatsApp contact id
status: str # in_progress / completed / abandoned
answers: dict # JSON: {"q1": "...", "q2": "..."}
organization_id: int
project_id: int
inserted_at / updated_at
is the actual source of truth for what the user answered, this is what admin dashboards/reports query, not checkpoint blobs.
Node taxonomy — retry/skip as session state, not a graph loop
- Session state carries a
no_answer_countper question plus the collected answers so far. - A conditional router (same shape as #1210
route_after_agent) decides, per node: re-ask the same question, advance to the next node, or apply the node's configuredon_exhaustedpolicy (skip /end_conversation/ escalate, open to adding more). - This keeps the retry/skip policy entirely config-driven and independent of the LLM/tool machinery.
Tool catalogue — same registry pattern as #1210, extended to write tools
- Reuse the AgentTool-style registry (name, description, args model, executor) built for the read-only agent. New entries:
recurring_reminder,one_off_reminder,delete_reminder, etc. - Tool selection is per-node, following OCS: each
llm_with_tools-typenode in the flow config carries its own"tools": [...] allowlist, not one list for the whole bot. This lets different steps in the same flow expose different tools (e.g. a reminder step offersrecurring_reminder/one_off_reminder, a later step offersdelete_reminderonly). The node's allowlist is filtered from the global registry and passed to the model for that turn, same as #1210 filters its fixed tool list.
Scheduled actions — Celery, decoupled from the live graph
- A reminder tool call writes a row (
fire_at, recurrence,session_id, payload) and returns immediately, it never blocks inside the conversation graph. - A Celery periodic task (existing app/celery/ convention) polls due actions and either sends a message directly (e.g. a WhatsApp nudge via Glific) or resumes the conversation graph for that session.
one_off_reminderfires once;recurring_reminderrecomputes its ownnext_fire_atafter firing;delete_remindercancels a pending row.
State persistence — LangGraph checkpointer (diverges from OCS)
- Unlike #1210's stateless single-request agent, this bot's conversations span many turns over potentially days. Reuse the
langgraph-checkpoint-postgrespattern already proven out in services/evaluations/iteration_checkpointer.py. - Each inbound message (Glific/WhatsApp webhook) re-invokes the graph on the same
thread_id (= conversation/session id); LangGraph reloads state from Postgres. interrupt() is used the same way the eval-iteration loop uses it, to pause for the user's next reply across an arbitrarily long gap.
Versioning
FlowConfigVersiongives us working-vs-published for free, the same wayConfigVersionalready does for/llm/call, an admin can edit a flow without affecting conversations already in progress on the previously published version.
High-level shape
Admin UI ──► FlowConfigVersion (nodes + edges + per-node tool allowlist, JSON)
│
▼
Graph builder (node "type" → Python class)
│
▼
StateGraph, compiled with a Postgres checkpointer, thread_id = session id
│
┌─────────────┼───────────────────┐
▼ ▼ ▼
question_node llm_with_tools_node router_node
(retry/skip (tools = this node's (branches on
count in state) own selected list) session state)
│
▼
Tool registry (recurring_reminder, one_off_reminder, delete_reminder, ...)
│
▼
Celery periodic task (scheduled_action table)
│
▼
Glific/WhatsApp delivery (resume graph OR send a message)
- 主要语言
- Python
- 星标
- 18
- 派生
- 10
- 平均合并
- 3 天 20 小时
- 30 天内合并 PR
- 14
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
ProjectTech4DevAI/kaapi-backend 的其他 Issue
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
ProjectTech4DevAI/kaapi-backend#889 ·
维护者通常 2 天内回复
-
enhancement
难度 2/5 1-3 小时 新手友好度 68/100
ProjectTech4DevAI/kaapi-backend#269 · 1 条评论 ·
维护者通常 2 天内回复
-
Penetration Testing: Integrate Shannon for reports可能已有人在做 @Prajna1999 今天认领。 未关闭
ProjectTech4DevAI/kaapi-backend#1232 · 已指派 1 人 ·
维护者通常 2 天内回复
-
Evaluation: TTS metrics for projects可能已有人在做 @Prajna1999 于 1 天前认领。 未关闭
ProjectTech4DevAI/kaapi-backend#1229 · 已指派 1 人 ·
维护者通常 2 天内回复
-
TTS Evaluation: Add more models可能已有人在做 @Prajna1999 于 1 天前认领。 未关闭
ProjectTech4DevAI/kaapi-backend#1228 · 已指派 1 人 ·
维护者通常 2 天内回复
查看 ProjectTech4DevAI/kaapi-backend 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 83/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 86/100
FuRongJun-1999/dsh-memory#65 ·
维护者通常 1 天内回复
-
ci needs-ac
难度 2/5 1-3 小时 新手友好度 75/100
Ikalus1988/MisakaNet#2930 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 78/100
Qiskit/qiskit-aer#2466 ·
-
area/cli
难度 2/5 1-3 小时 新手友好度 82/100