performance issues due to breadth first execution of grapqhl queries in case of async resolvers during calls burst
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Anfängerfreundlichkeit
- 25/100
Rechercherichtung
Beginne mit dem asyncio-Proof-of-Concept des Issues und dem Einstiegspunkt schema.execute_async und verfolge anschließend, wie asynchrone GraphQL-Felder geplant und aufgelöst werden. Vergleiche dieses Verhalten mit der Task-Warteschlange von asyncio bei nebenläufigen Abfragen. Als abgeschlossen gilt die Aufgabe, wenn ein abgestimmter Scheduling- oder Prioritätsansatz vorliegt, der anhand des gemeldeten Burst-Szenarios demonstriert wird, ohne die Ausführungszeit oder den Speicherverbrauch zu verschlechtern.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
When we receive calls burst we found that the calls "wait each other" (i.e. the first call of the burst waits the last one).
This result in degradation of performances in both time of execution and memory consumption in the server because we have to keep many calls in fly.
This is particularly evident in big graphql queries where the users request many fields and we have several depth in the queries where each level has many async fields.
Just to be TLDR, looking the implementations of graphql and asyncio we understood that this is due to the following:
- graphql breadth first way to schedule and resolve the fields
- asyncio internal FIFO queue of tasks to be executed
As an example lets have queries like that, where data may be like beer vendors and we want for each beer vendor many fields that describes that vendor, a1...a100, b1...b100, ...:
query {
data {
a1 {
b1 {
c1
...
c100
}
...
b100 {
c1
...
c100
}
}
...
a100 { ... }
}
}
If we have n of this calls coming in burst when we arrive to the depth of the c fields we have many many task scheduled in the asyncio queue.
If we check the of order of execution we have that the first query, on each level, "waits" the other queries, because all the queries schedules a lot of tasks.
In the proof of concept, that you may find at the end of the post, you can verify the order of execution of the resolvers.
It could be very nice to have some sort of priority in the order to let the first query not wait the scheduling and resolve of all the queries before ending.
I understand that this is something between graphql and asyncio but i think it could affect the use of graphql in environments receiving many calls.
Fixes, helps and hints in how to improve this would be very appreciated.
import asyncio
from graphene import ObjectType, Schema, String, Field
FIELD_NUMBER = 2
CONCURRENT_QUERIES = 10
def make_resolver(i, j=None):
async def resolver(self, info):
print(f"START query {info.context['query_number']} | a{i} | b{j}")
await asyncio.sleep(0.001)
print(f"END query {info.context['query_number']} | a{i} | b{j}")
return i
return resolver
def create_fields():
fields = {}
for i in range(FIELD_NUMBER):
inner_fields = {}
for j in range(FIELD_NUMBER):
inner_fields[f"b{j}"] = String()
inner_fields[f"resolve_b{j}"] = make_resolver(i, j)
MyType = type(
f"MyType",
(ObjectType,),
inner_fields,
)
fields[f"a{i}"] = Field(MyType)
fields[f"resolve_a{i}"] = make_resolver(i)
return fields
async def make_query(schema, query_number):
inner_query_values = [f"b{i}" for i in range(FIELD_NUMBER)]
query_values = [
"a%s {%s}" % (i, " ".join(inner_query_values)) for i in range(FIELD_NUMBER)
]
query_string = "{ %s }" % (" ".join(query_values),)
await schema.execute_async(
query_string, context_value=dict(query_number=query_number)
)
async def main():
Query = type("Query", (ObjectType,), create_fields())
schema = Schema(query=Query)
await asyncio.gather(*[make_query(schema, i) for i in range(CONCURRENT_QUERIES)])
asyncio.run(main())
- Vorherrschende Sprache
- Python
- Sterne
- 531
- Forks
- 147
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus graphql-python/graphql-core
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 50/100
graphql-python/graphql-core#272 · 1 Kommentar ·
-
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 55/100
graphql-python/graphql-core#269 · 1 Kommentar ·
-
Publish a major version Offen
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
graphql-python/graphql-core#267 · 1 Kommentar ·
-
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 45/100
graphql-python/graphql-core#257 ·
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 25/100
graphql-python/graphql-core#247 · 8 Kommentare ·
Alle Issues in graphql-python/graphql-core
Ähnliche Issues
-
agent-ready documentation needs-triage
Schwierigkeit 1/5 1-3 Stunden Anfängerfreundlichkeit 88/100
-
documentation
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 91/100
-
workflow-status page template still says reusable workflows are "triggered only by workflow_call:" Offen
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 92/100
-
instance instance add
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 72/100
searxng/searx-instances#939 · 1 Kommentar ·
-
area-deployment area-integrations triage:bot-seen
Schwierigkeit 2/5 Ein halber Tag Anfängerfreundlichkeit 86/100