context length and trajectory problems
Los mantenedores suelen responder en 6 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 30/100
Línea de trabajo
Empieza por rastrear cómo se crea y se consume task.md durante el rollout, cómo se ensamblan los prompts durante la reflexión y dónde se genera trace raw.txt. Compara las rutas Claude Code Exec y skillopt sleep para determinar si se conservan el contexto, las muestras fallidas y las trayectorias intermedias de las herramientas; se considerará terminado cuando el contexto y los datos de trayectoria solicitados lleguen a la reflexión sin truncarse.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Context size issue: For example, when training on the SearchQA dataset using Claude Code Exec as the backend, the context and question are written into a task.md file, which Claude reads to generate answers. This avoids the problem of input text being too long during the rollout phase, making it impossible for Claude to fully process the content. However, during the reflection phase, Skillopt currently does not support full Claude Code Exec as the backend. Instead, relevant information—including the context needed to answer the question—must be included directly in the prompt, then calling claude once. This leads to potential context truncation issues.
Alse skillopt sleep:During the rollout phase, skillopt sleep did not use full Claude execution but instead treated it as a chat endpoint, so context was likely limited. In the reflection phase, reference materials for context probably weren't fed into the reflector either, and each failed sample's question, answer, and failure reason were truncated.
Do you have any plans to optimize the above two context-related scenarios in the future?
The issue regarding the intermediate process trajectory: Taking the training data of the Searchqa dataset as an example, when using the Claude code execution backend, the intermediate execution trajectory of Claude code (such as the intermediate thinking process, tool calls, etc.) was not parsed and saved. Although I noticed that a trace raw.txt file was generated in the code, the content was just a very simple summary. Additionally, the trajectory process was not sent to the reflection stage.
Does the official have any optimization plans for this issue? If this situation can be supported, then the dataset only needs to provide the questions and answers. The agent will provide the intermediate trajectory and the final result, and all of them will be sent to the reflect stage. The reflector can simultaneously analyze the agent's output and the intermediate process, and propose more targeted skill optimization suggestions.
- Lenguaje dominante
- Python
- Estrellas
- 18k
- Forks
- 1.7k
- Merge medio
- 6 d 18 h
- PR fusionados (30 d)
- 12
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/SkillOpt
-
Same unconfined `predictions/<item_id>` path pattern remains in four benchmark rolloutsPosiblemente ocupada @adongwanai la tomó hace 2 días. Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
Los mantenedores suelen responder en 6 días
-
项目还在迭代嘛?Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 10/100
microsoft/SkillOpt#291 · 1 comentario ·
Los mantenedores suelen responder en 6 días
-
Release cut for the Aug adopt/webui hardening + 2 residual staging gapsPosiblemente ocupada @RohithPariki la tomó hace 19 días. Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 42/100
microsoft/SkillOpt#288 · 2 comentarios ·
Los mantenedores suelen responder en 6 días
-
skillopt-sleep Codex harvest ingests its own headless replay sessionsPosiblemente ocupada @kaluli123123 la tomó hace 20 días. Abierto
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
microsoft/SkillOpt#286 · 3 comentarios ·
Los mantenedores suelen responder en 6 días
-
Proposal: optional offline typed-decision judge for rubric scoring (SemIf / NanoJev pattern)Abierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 28/100
microsoft/SkillOpt#283 · 1 comentario ·
Los mantenedores suelen responder en 6 días
Todos los issues de microsoft/SkillOpt
Issues similares
-
Add `django-upgrade` to the CIAbiertodependencies feature github_actions good first issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
wemake-services/wemake-django-template#3149 ·
Los mantenedores suelen responder en 1 día
-
[request] vsg/1.1.16Abiertoupstream update
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
conan-io/conan-center-index#31142 ·
Los mantenedores suelen responder en 1 día
-
area:core bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
request-theme
Dificultad 2/5 Menos de una hora Aptitud para principiantes 70/100
LizardByte/ThemerrDB#8877 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
area/install-update comp/gateway P0 sweeper:risk-compatibility type/bug
Dificultad 2/5 Menos de una hora Aptitud para principiantes 72/100
NousResearch/hermes-agent#135997 · 3 comentarios ·
Los mantenedores suelen responder en 1 día