ANTLR grammar fails to parse graph queries
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
Línea de trabajo
Start with grammar/Kql.g4 and reproduce the issue by generating the Python parser with ANTLR 4.13.2 using the command in the report. Check the documented graph-match, graph-shortest-paths, make-graph, and range queries listed in the issue. Done means all six grammar gaps are fixed, generation is clean, and the listed queries parse without changing the existing 1,285 documentation-query results.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Feature request to complete support for graph-match, make-graph, etc. in the KQL ANTLR grammar grammar/Kql.g4.
AI analysis of the syntax support gap:
Summary
The ANTLR reference grammar in grammar/ cannot parse most of the documented
graph-semantics queries, nor a numeric range written without spaces
(between (1..3), -[e*1..3]->). All of the queries below are taken from, or
follow the syntax of, the Microsoft Learn KQL reference, and Kusto accepts them.
The hand-written parser in src/Kusto.Language is not affected; this is about
the .g4 files only.
Found while generating a Python parser from the grammar with ANTLR 4.13.2
(java -jar antlr-4.13.2-complete.jar -Dlanguage=Python3 -visitor Kql.g4).
1. A graph pattern is parsed as one element per comma
graphMatchOperator:
GRAPHMATCH
(Parameters+=relaxedQueryOperatorParameter)*
Patterns+=graphMatchPattern (',' Patterns+=graphMatchPattern)*
...
graphMatchPattern:
Node=graphMatchPatternNode
| UnnamedEdge=graphMatchPatternUnnamedEdge
| NamedEdge=graphMatchPatternNamedEdge;
Each comma-separated pattern is a single node or edge, so any pattern with an
edge in it fails. graph-shortest-paths reuses the rule and fails the same way.
let E = datatable(s:string, t:string)["A","B"];
E | make-graph s --> t with_node_id=id
| graph-match (a)-[e]->(b) project a.id, b.id
line 3:16 mismatched input '-[' expecting {<EOF>, ';'}
Documented syntax: graph-match operator,
"Graph pattern notation" — a pattern is a sequence of nodes joined by edges,
and several such sequences may be separated by commas.
Suggested fix:
graphMatchPattern:
Nodes+=graphMatchPatternNode (Edges+=graphMatchPatternEdge Nodes+=graphMatchPatternNode)*;
graphMatchPatternEdge:
UnnamedEdge=graphMatchPatternUnnamedEdge
| NamedEdge=graphMatchPatternNamedEdge;
2. Anonymous nodes and anonymous variable-length edges are rejected
graphMatchPatternNode and graphMatchPatternNamedEdge both require a name,
but the notation table documents () for an anonymous node and
-[*3..5]- for an anonymous variable-length edge.
E | make-graph s --> t with_node_id=id
| graph-match (a)-->()-[*1..3]->(b) project a.id, b.id
Suggested fix: make both names optional.
graphMatchPatternNode:
'(' (Name=identifierOrKeywordOrEscapedName)? ')';
graphMatchPatternNamedEdge:
OpenBracket=(DASH_OPENBRACKET | LESSTHAN_DASH_OPENBRACKET)
(Name=identifierOrKeywordOrEscapedName)?
(Range=graphMatchPatternRange)?
CloseBracket=(CLOSEBRACKET_DASH_GREATERTHAN | CLOSEBRACKET_DASH)
;
3. 1..3 lexes as the real 1. followed by .3
NonIntegerNumber accepts a trailing-dot real (1.), and the lexer's longest
match prefers it to the integer 1. So 1..3 becomes 1. then .3, and both
the variable-length edge range and between fail whenever the range is written
without spaces:
range x from 1 to 3 step 1 | where x between (2..3)
no viable alternative at input '.3'
E | make-graph s --> t with_node_id=id
| graph-match (a)-[p*1..3]->(b) project a.id, b.id
The graph documentation writes every range this way (-[e*1..5]-,
-[reports*1..5]-), so this blocks most of its examples even after fix 1.
Suggested fix: a trailing-dot real must not be followed by a second dot.
ANTLR has no target-neutral negative lookahead, so this needs a semantic
predicate in each target's language; for the Python target:
fragment NonIntegerNumber:
('0'..'9')+ '.' {self._input.LA(1) != 46}? ('0'..'9')* Exponent?
| ('0'..'9')+ Exponent
;
(46 is '.'. The C# and Java spelling is {_input.LA(1) != '.'}?.) 1.,
1.5, 1.5e3 and 1.e2 all still lex as reals.
4. make-graph accepts only one node table
makeGraphTablesAndKeysClause:
WITH Table=invocationExpression ON Column=simpleNameReference;
The documented syntax is with Nodes1 on NodeId1 [, Nodes2 on NodeId2]
(make-graph operator):
E | make-graph s --> t with People on pid, Companies on cid
| graph-match (a)-->(b) project a, b
mismatched input ',' expecting {<EOF>, ';'}
Suggested fix:
makeGraphTablesAndKeysClause:
WITH Tables+=makeGraphTableAndKey (',' Tables+=makeGraphTableAndKey)?;
makeGraphTableAndKey:
Table=invocationExpression ON Column=simpleNameReference;
5. partitioned-by requires a dotted path
makeGraphPartitionedByClause:
PARTITIONEDBY Entity=entityPathOrElementExpression '(' SubQuery=contextualSubExpression ')';
entityPathOrElementExpression needs at least one . or [...], but the
documented form is a plain column, partitioned-by PartitionColumn (GraphOperator):
E | make-graph s --> t with N on id partitioned-by tenant (graph-match (a)-->(b) project a.id)
mismatched input '(' expecting {'.', '['}
Suggested fix: PARTITIONEDBY Column=simpleNameReference '(' ... ')'.
6. with_node_id= and output= are not accepted as operator parameters
with_node_id and output are keyword tokens (WITH_NODE_ID, OUTPUT), and
relaxedQueryOperatorParameter lists neither, so these documented forms fail:
E | make-graph s --> t with_node_id=id | graph-to-table nodes with_node_id=NodeId
E | make-graph s --> t with_node_id=id
| graph-shortest-paths output=all (a)-[p*1..3]->(b) project a.id, b.id
Suggested fix: add WITH_NODE_ID and OUTPUT to the NameToken
alternatives of relaxedQueryOperatorParameter.
Verification
With all six fixes, ANTLR generation is clean, every query above parses, and
the 1,285 documentation queries we already parsed produce identical parse
results. Re-harvesting the KQL reference documentation then also parses the
code blocks on the graph-match, graph-shortest-paths, make-graph and graph
function pages, which were previously all rejected.
- Lenguaje dominante
- C#
- Estrellas
- 681
- Forks
- 122
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/Kusto-Query-Language
-
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
-
Dificultad 3/5 1-2 días Aptitud para principiantes 68/100
-
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
microsoft/Kusto-Query-Language#194 · 3 comentarios ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 52/100
microsoft/Kusto-Query-Language#189 · 1 reacción ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 48/100
microsoft/Kusto-Query-Language#188 · 1 comentario ·
Todos los issues de microsoft/Kusto-Query-Language
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
AvaloniaUI/Avalonia#22420 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 90/100
dotnet/SqlClient#4823 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
type/automation type/tech-debt
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
clockworklabs/SpacetimeDB#6124 ·
Los mantenedores suelen responder en 1 día
-
area-integrations
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Los mantenedores suelen responder en 1 día