Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

Python: add PyMongo read results as sources for `py/sql-injection` in second-order SQL construction flows

Aberta
#21,775 1 comentário 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
4/5
Tempo estimado
3-5 dias
Facilidade para iniciantes
48/100
Tipo de issue
Funcionalidade
Clareza
Razoavelmente clara
Status de atividade
Pouca atividade
Stack de tecnologia
python, sql
Domínio
databases, security

Direção de pesquisa

Comece lendo PyMongo.qll, PEP249.qll e a modelagem existente de sources e sinks de py/sql-injection. Rastreie como os resultados de find, find_one e find_one_and_* poderiam entrar no fluxo existente de SQL-injection e, em seguida, valide o comportamento pretendido com um exemplo reduzido de segundo estágio e os testes da query; o trabalho estará concluído quando os valores persistidos do PyMongo chegarem ao sink execute existente sem uma nova modelagem de sinks.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

py/sql-injection already appears to model the sink side correctly through the existing DB-API / PEP249.qll coverage for execute(...). The gap seems to be on the source side for a common second-order pattern: values read from MongoDB with PyMongo are later reused in dynamically constructed SQL. I ran into this while triaging KBase Metrics (CVE-2022-4860), but the underlying issue is broader than that one project.

A reduced example looks like this:

from pymongo import MongoClient
import psycopg2

def sync_users():
    users = []
    for record in MongoClient(uri).auth.users.find({"role": "dev"}, {"user": 1, "_id": 0}):
        users.append(record["user"])

    in_clause = "', '".join(users)
    sql = (
        "update user_info set active = true "
        "where username in ('" + in_clause + "')"
    )

    cur = psycopg2.connect(dsn).cursor()
    cur.execute(sql)

My reading of the current modeling is that this flow falls between two existing pieces: PyMongo.qll models collection operations for NoSQL semantics, while py/sql-injection starts from active threat-model sources that do not seem to cover data read back from PyMongo collections. As a result, the query has the right sink and the right string-building path shape, but no source that can reach it.

I do not think this needs a new query or wider sink modeling. The fix seems fairly contained: add source coverage for values obtained from common PyMongo read APIs such as find, find_one, and find_one_and_*, so that those results can participate in the existing py/sql-injection flow. If widening default behavior is a concern, this could also live behind an opt-in threat-model bucket for persisted database results rather than being treated as generic local input.

This pattern is common in real Python codebases, especially in cron jobs, reporting jobs, migration scripts, sync workers, and ETL-style code that bridges Mongo-backed application state into relational stores. Bandit's B608 already flags the same family syntactically by recognizing SQL-shaped string construction passed to execute(), so there is at least external evidence that this is a practical and recurring pattern. CodeQL seems close to covering it already; the missing piece is verifiable semantic coverage for PyMongo-backed persisted data flowing into the existing SQL sinks.

Linguagem predominante
CodeQL
Estrelas
10.1k
Forks
2.1k
Merge médio
2d 16h
PRs com merge (30d)
143

Guia de contribuição

Abrir o guia de contribuição

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de github/codeql

Todas as issues de github/codeql

Issues semelhantes

Mais issues de Databases

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.