[bot] Cohere: Embed Jobs API (`client.embed_jobs.create()`) not instrumented
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bien especificado
- Estado de actividad
- Tranquilo
- Stack tecnológico
- python
- Área
- observability-sre
Línea de trabajo
Comienza con py/src/braintrust/integrations/cohere/patchers.py e integration.py y luego inspecciona las pruebas existentes en py/src/braintrust/integrations/cohere/test_cohere.py. Añade cobertura para los métodos de Cohere Embed Jobs y sus equivalentes async, siguiendo los patrones de instrumentación existentes. Se considera terminado cuando la creación del job registra el dataset y la configuración del modelo solicitados, además del ID del job y el estado inicial, con pruebas que cubran los nuevos spans.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
The Cohere Python SDK's Embed Jobs API — client.embed_jobs.create() / client.embed_jobs.get() / client.embed_jobs.list() (and the AsyncClient equivalents) — is not instrumented. This is Cohere's bulk/batch embedding execution surface: it launches an async job that reads a Dataset of type embed-input, runs the configured embedding model (model, input_type, embedding_types) over every record, and writes the resulting vectors to a new embed-output Dataset. It is the async, large-scale counterpart to the synchronous client.embed() call, which is instrumented.
Calls to client.embed_jobs.create() fall through uninstrumented today — the Cohere integration only patches BaseCohere/AsyncBaseCohere (chat, chat_stream, embed, rerank) and the v2 client and audio transcription client. No patcher touches the embed_jobs resource, so job creation produces zero Braintrust tracing (no span for the submitted job, its model/config, or its terminal status).
What is missing
| Cohere method | Instrumented? |
|---|---|
client.chat() / client.chat_stream() (v1 + v2) |
Yes |
client.embed() (v1 + v2) |
Yes |
client.rerank() (v1 + v2) |
Yes |
client.audio.transcriptions.create() |
Yes |
client.embed_jobs.create() |
No |
client.embed_jobs.get() |
No |
AsyncClient.embed_jobs.create() |
No |
At minimum, instrumentation should create a span for embed_jobs.create() capturing:
- Input: dataset ID, model,
input_type,embedding_types, job name/truncate settings - Output: job ID and initial status (
processing/complete/failed) - Metadata: model name, dataset reference
This mirrors the pattern already accepted for the analogous async-batch-execution gaps in sibling integrations: OpenAI Batch API (#295), Google GenAI Batch API (#332), and Mistral Batch Jobs API (#272) — this is the same class of gap (an async, job-based generative-execution API that bypasses the wrapper) applied to Cohere's bulk embeddings surface.
Braintrust docs status
not_found — The Cohere integration page documents only: chat completion spans (cohere.chat / cohere.chat_stream), tool call spans, embedding spans (cohere.embed), rerank spans (cohere.rerank), and audio transcription spans (cohere.audio.transcriptions.create). Embed Jobs, batch/bulk embeddings, the Classify API, and the v1 generate() endpoint are not mentioned anywhere on the page.
Upstream sources
- Cohere Embed Jobs API reference (
POST /v1/embed-jobs): https://docs.cohere.com/reference/create-embed-job - Cohere conceptual guide, "Batch Embedding Jobs with the Embed API": https://docs.cohere.com/v2/docs/embed-jobs-api
- Cohere Python SDK (
cohere-python) exposes this asclient.embed_jobs.create(...)/.get(...)/.list(...)on bothcohere.Clientandcohere.AsyncClient
Local files inspected
py/src/braintrust/integrations/cohere/patchers.py— definesChatPatcher,ChatStreamPatcher,EmbedPatcher,RerankPatcher(+ async/v2 variants) andTranscriptionsCreatePatcher/AsyncTranscriptionsCreatePatcher; noembed_jobspatcher of any kindpy/src/braintrust/integrations/cohere/integration.py—CohereIntegration.patchersregisters only the five patchers above;embed_jobsis absent from the registrypy/src/braintrust/integrations/cohere/tracing.py— noembed_jobwrapper functions existpy/src/braintrust/integrations/cohere/test_cohere.py— no test cases referenceembed_jobs- Repo-wide search for
embed_job/embedjob(case-insensitive) returns zero matches outside this issue
Relationship to existing issues
Distinct from #488 (Cohere v1 generate()/generate_stream()) and #343 (Cohere Classify API) — both already open and unrelated to Embed Jobs. Same class of gap as the already-filed provider Batch API issues (#295, #332, #272), applied to a surface (Cohere Embed Jobs) none of those cover.
- Lenguaje dominante
- Python
- Estrellas
- 20
- Forks
- 18
- Merge medio
- 21 h 8 min
- PR fusionados (30 d)
- 81
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de braintrustdata/braintrust-sdk-python
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
braintrustdata/braintrust-sdk-python#797 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
braintrustdata/braintrust-sdk-python#774 ·
Los mantenedores suelen responder en 1 día
-
Migrate legacy HTTPConnection callers to the policy-aware Transport incrementallyPosiblemente ocupada @AbhiPrasad la tomó hace 1 día. Abiertopython
braintrustdata/braintrust-sdk-python#839 · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
new-integration
Dificultad 4/5 3-5 días Aptitud para principiantes 68/100
braintrustdata/braintrust-sdk-python#808 ·
Los mantenedores suelen responder en 1 día
-
new-integration
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
braintrustdata/braintrust-sdk-python#807 ·
Los mantenedores suelen responder en 1 día
Todos los issues de braintrustdata/braintrust-sdk-python
Issues similares
-
#bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
apache/superset#44923 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
lawndoc/stack-back#123 ·
-
Add: EntuneAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
AbdelStark/awesome-typesafe-jev#187 ·
Los mantenedores suelen responder en 1 día
-
bug good first issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
repowise-dev/repowise#2966 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 2 días