Python: extract_range() drops multi-call assistant message when filtering out just one of its parallel tool results
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 3/5
- Tiempo estimado
- 1-2 días
- Aptitud para principiantes
- 78/100
Línea de trabajo
Start in python/semantic_kernel/contents/history_reducer/chat_history_reducer_utils.py by reading extract_range() and get_call_result_pairs(), then run the parallel-call reproduction from the issue. Update the pair association so one assistant call index accounts for all its result indices; done means filtering one result preserves the call and its other result without orphaning either message.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Describe the bug
extract_range() in python/semantic_kernel/contents/history_reducer/chat_history_reducer_utils.py is used by ChatHistorySummarizationReducer / ChatHistoryTruncationReducer (via preserve_pairs=True) to make sure a function call is never separated from its result when the history is reduced.
The pair lookup is built like this (current main):
pair_map = {}
if preserve_pairs:
pairs = get_call_result_pairs(history)
for cidx, ridx in pairs:
pair_map[cidx] = ridx
pair_map[ridx] = cidx
get_call_result_pairs() returns one (call_index, result_index) tuple per function call, so when a single assistant message contains more than one FunctionCallContent (the normal shape for parallel/multi tool calling), several pairs share the same call_index. Because pair_map is a plain dict, pair_map[cidx] = ridx silently overwrites earlier entries, so only the last result for that call index survives in the forward direction. The reverse entries (pair_map[ridx] = cidx) don't collide, so each result still points back at the call correctly - only the call's own forward pointer is wrong.
Later in extract_range, when the call message's own turn comes up, its keep/skip decision (and the filter_func "skip both" check) is made using that single, possibly-wrong paired_idx, not the full set of results tied to that call. If a filter_func is used to drop just one of several parallel tool results, the call message ends up being evaluated only against the result that happens to still be in pair_map[cidx], which can cause the call to be dropped even though another of its own results is being kept - leaving a TOOL result message in the output with no corresponding assistant tool_calls message at all.
To Reproduce
Ran this against the actual installed semantic-kernel package (pip show semantic-kernel -> 1.44.1; diffed byte-for-byte identical to the current main copy of chat_history_reducer_utils.py before running):
from semantic_kernel.contents.chat_message_content import ChatMessageContent
from semantic_kernel.contents.function_call_content import FunctionCallContent
from semantic_kernel.contents.function_result_content import FunctionResultContent
from semantic_kernel.contents.utils.author_role import AuthorRole
from semantic_kernel.contents.history_reducer.chat_history_reducer_utils import extract_range, get_call_result_pairs
msg0 = ChatMessageContent(role=AuthorRole.USER, content="weather?")
msg1 = ChatMessageContent(
role=AuthorRole.ASSISTANT,
items=[
FunctionCallContent(id="call_paris", name="get_weather", arguments='{"city":"Paris"}'),
FunctionCallContent(id="call_tokyo", name="get_weather", arguments='{"city":"Tokyo"}'),
],
)
msg2 = ChatMessageContent(role=AuthorRole.TOOL, items=[FunctionResultContent(id="call_paris", name="get_weather", result="15C")])
msg3 = ChatMessageContent(role=AuthorRole.TOOL, items=[FunctionResultContent(id="call_tokyo", name="get_weather", result="22C")])
msg4 = ChatMessageContent(role=AuthorRole.ASSISTANT, content="done")
history = [msg0, msg1, msg2, msg3, msg4]
print(get_call_result_pairs(history))
# [(1, 2), (1, 3)] <- both pairs share call_index 1
filter_func = lambda m: m is msg3 # drop only the Tokyo result, e.g. a moderation/error filter
out = extract_range(history, start=0, end=len(history), filter_func=filter_func, preserve_pairs=True)
for m in out:
print(m.role, [getattr(it, "id", None) for it in m.items] if m.items else m.content)
Output:
AuthorRole.USER weather?
AuthorRole.TOOL ['call_paris']
AuthorRole.ASSISTANT done
The assistant message with the two tool_calls (msg1) is gone entirely, even though call_paris's own result (msg2) is kept right after it with no call message in front of it.
Expected behavior
Filtering out only call_tokyo's result should not affect call_paris's pair. Expected output keeps the call message and the Paris result together, and drops only the Tokyo result:
AuthorRole.USER weather?
AuthorRole.ASSISTANT [call_paris, call_tokyo] (or with call_tokyo's FunctionCallContent stripped)
AuthorRole.TOOL ['call_paris']
AuthorRole.ASSISTANT done
At minimum, the call message and call_paris's result should never both disappear/appear inconsistently relative to each other - right now the call is dropped solely because of an unrelated sibling call's filtered result.
Platform
- Language: Python
- Source: pip package
semantic-kernel==1.44.1, and confirmed the relevant file is unchanged on the currentmainbranch - File:
python/semantic_kernel/contents/history_reducer/chat_history_reducer_utils.py, functionextract_range()(thepair_mapconstruction) together withget_call_result_pairs()
Additional context
This is a distinct root cause from the interleaved-messages reordering bug already reported and fixed in the still-open #14165 (Fix extract_range reordering messages when preserving function call/result pairs). I checked that PR's rewritten implementation and it keeps the exact same pair_map[cidx] = ridx / pair_map[ridx] = cidx construction, so this dict-collision bug reproduces identically against that PR's version of the function too - it is not fixed as a side effect of #14165 and needs its own fix in how pair_map associates a call index with all of its result indices (e.g. pair_map[cidx] holding a list of result indices instead of a single int).
Related but different: #12708 and #13062 (both closed by the stale bot, not fixed) describe tool-call/result pairs being orphaned when a pair straddles the [start, end) boundary. This report is about pairs being mishandled even when everything is fully inside the requested range, purely because more than one result shares a call index.
- Lenguaje dominante
- C#
- Estrellas
- 28.6k
- Forks
- 4.8k
- Merge medio
- 12 h 20 min
- PR fusionados (30 d)
- 12
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/semantic-kernel
-
python triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
microsoft/semantic-kernel#14491 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
python triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
microsoft/semantic-kernel#14490 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Python: [Python] structured_outputs_transform reuses ChatHistory across calls (prompt pollution)Abiertopython triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
microsoft/semantic-kernel#14483 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
.NET python triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
microsoft/semantic-kernel#14482 · 2 comentarios ·
Los mantenedores suelen responder en 2 días
-
Python: [Python] as_agent_framework_tool drops parameter defaults (optionals become required)Abiertopython triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
microsoft/semantic-kernel#14481 ·
Los mantenedores suelen responder en 2 días
Todos los issues de microsoft/semantic-kernel
Issues similares
-
0 - Backlog Bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
BrighterCommand/Brighter#4444 ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
area:frontend bug FE P3
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
klasolsson81/jobbliggaren#1915 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
microsoft/vscode-copilotstudio#431 ·
Los mantenedores suelen responder en 2 días