Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Resolving `MissingGreenlet` error when using async drivers.

未关闭
#419 3 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
25/100
Issue 类型
缺陷
描述清晰度
需要澄清
活跃度
停滞
技术栈
graphql, postgresql, python, sqlalchemy
领域
api, backend, databases

调研方向

从 graphene_sqlalchemy/batching.py 开始,重点查看链接行中的 RelationshipLoader.batch_load_fn,并使用所述的 aiohttp、asyncpg、PostgreSQL 和 GraphQL 配置复现 MissingGreenlet 错误。将 SQLAlchemy 2 路径与提供的 greenlet_spawn workaround 进行比较;完成的标准是达成一个一致认可的通用修复方案,或清楚记录受支持的异步驱动行为。

由索引模型根据 Issue 内容生成。

描述

I'm not sure if other people have run into this problem - I searched for days and couldn't find a clear answer. But eventually I figured out a solution, so wanted to document it here in case other people do run into it.

Background

Our setup uses a Postgres database which provides data for a GraphQL API, for which we initially used Flask, following the example. However, our schema uses a lot of 1:N relationships, so our performance was degraded due to the N+1 Round Trip Problem. The fix for this is, supposedly, very simple: just set batching=True in the meta class. However, doing this triggers the use of dataloaders, which operate asynchronously. So we switched from Flask to using the aiohttp web framework and asyncpg database driver, which are better suited for asynchronous tasks. This is when we first stumbled on the MissingGreenlet error: sqlalchemy.exc.MissingGreenlet: greenlet_spawn has not been called; can't call await_only() here. Was IO attempted in an unexpected place?

Issue

The problem here comes from the RelationshipLoader.batch_load_fn() (link). As the function comment explains, that function has to "piggyback on some internal APIs" of sqlalchemy. As a result (I believe), the necessary greenlet_spawn function never gets called, and the above error is thrown. (Greenlets, AFAICT, are small thread-like objects that allow synchronous requests to be performed in an asynchronous environment. In this situation, the dataloader is what actually materializes the expected attributes, so it is necessary to make these calls synchronously)

Solution

So on to the fix: we just need to call greenlet_spawn ourselves. This can be done by simply redefining the batch_load_fn attribute on the RelationshipLoader class, so that all invocations call the corrected function rather than the incorrect code. My stripped down code is below (I removed some extraneous comments/checks/conditions from batch_load_fn for simplicitly, make sure that you compare to the original code and update the correct path for your situation):

from graphene_sqlalchemy.batching import RelationshipLoader
from sqlalchemy.orm import Session
from sqlalchemy.util import immutabledict
from sqlalchemy.util._concurrency_py3k import greenlet_spawn


def patch_relationship_loader() -> None:
    """
CALL THIS FUNCTION ONCE AS PART OF SETUP CODE
    """
    RelationshipLoader.batch_load_fn = batch_load_fn


async def batch_load_fn(self: RelationshipLoader, parents: Any) -> list[Any]:
    """
FOLLOWS THE `SQL_VERSION_HIGHER_EQUAL_THAN_2` PATH
    """
    child_mapper = self.relationship_prop.mapper
    parent_mapper = self.relationship_prop.parent
    session = Session.object_session(parents[0])

    for parent in parents:
        assert session is Session.object_session(parent)
        assert session and parent not in session.dirty

    states = [(sqlalchemy.inspect(parent), True) for parent in parents]

    query_context = None
    if session:
        parent_mapper_query = session.query(parent_mapper.entity)
        query_context = parent_mapper_query._compile_context()
# CALL greenlet_spawn HERE RATHER THAN CALLING _load_for_path DIRECTLY
        await greenlet_spawn(
            self.selectin_loader._load_for_path,
            query_context,
            parent_mapper._path_registry,
            states,
            None,
            child_mapper,
            None,
            None,  # recursion depth can be none
            immutabledict(),  # default value for selectinload->lazyload
        )
    result = [
        getattr(parent, self.relationship_prop.key) for parent in parents
    ]
    return result

Conclusion

I don't have enough knowledge of graphene-sqlalchemy or other database drivers to know what a universal solution would look like (or even if this is a universal problem - the dearth of information about this issue suggests not), so I don't want to submit a PR for this change, but at least in our situation this was the best fix. Hope it helps someone else out there!

主要语言
Python
星标
985
派生
224
PR 合并指标
30 天内没有已合并 PR

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

graphql-python/graphene-sqlalchemy 的其他 Issue

查看 graphql-python/graphene-sqlalchemy 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。