Performance idea: Partition before executing current uniform indentation search
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 20/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Necesita aclaración
- Estado de actividad
- Estancado
- Stack tecnológico
- ruby
- Área
- performance, tooling
Línea de trabajo
Empieza con la prueba que falla en la rama schneems/partition y sigue las llamadas a Ripper.parse del algoritmo de búsqueda actual. Lee cómo se gestionan la indentación y los pares kw/end, y después compara la partición o los pasos de expansión más grandes con el comportamiento existente. Se considera terminado cuando el caso de nueve mil líneas evita el timeout sin perder la calidad del resultado del error de sintaxis.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
This is a failing test: https://github.com/zombocom/dead_end/tree/schneems/partition. The file is nine thousand lines and it takes a tad over 1 second to parse which means it hits the timeout.
We are already fairly well optimized for the current algorithm so to be able to handle arbitrarily large-sized files we will need a different strategy.
The current algorithm takes relatively small steps in the interest in producing a good end result. That takes a long time.
Here's my general idea: We can split up the file into multiple large chunks before running the current fine-grained algorithm. At a high level: split up the file into 2 parts and see which holds the syntax error. If we can isolate the problem to only half the file then we've dropped processing time in half (relatively). We can run this partition step a few times.
The catch is that some files (such as the one in the failing test cannot be split without introducing a syntax error (since it starts with a class declaration and ends with an end). To account for this we will need to split in a way that's lexically aware.
For example on that file, I think the algorithm would determine that it can't do much with indentation 0 so it would have to go to the next indentation, there it could see there are N chunks of kw/end pairs, it could divide into N/2 and see if one of those sections holds all of the syntax errors. We could perform this division several times to arrive at a subset of the larger problem, then run the original search array on it.
The challenge is, that we will essentially need to build an inverse of the existing algorithm. Instead of starting with a single line and expanding towards indentation zero, we'll start with all the lines and reduce towards indentation max.
The expensive part is checking code is valid via Ripper.parse, sub dividing large files into smaller files can help us isolate problems sections with fewer parse calls, but we've got to make sure the results are as good.
An alternative idea would be to use the existing search/expansion logic to perform more expansions until a set of N blocks are generated then check all of them at once. Then once the document problem is isolated, go back and re-parse only the N blocks with the existing. Algorithm. (Basically the same idea as partitioning, but we're working from the same direction as the current algorithm, just taking larger steps (which means fewer Ripper.parse) calls. However we would still need a way to sub-divide the blocks with this process in the terminal case that the syntax error is on indentation zero and the document is massive and all within one kw/end pair.
- Lenguaje dominante
- Ruby
- Estrellas
- 350
- Forks
- 17
- Merge medio
- 48 min
- PR fusionados (30 d)
- 5
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de ruby/syntax_suggest
-
Dificultad 4/5 3-5 días Aptitud para principiantes 38/100
ruby/syntax_suggest#258 · 6 comentarios ·
-
Accidental if instead of a block Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 30/100
ruby/syntax_suggest#206 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 42/100
ruby/syntax_suggest#205 · 1 comentario ·
-
RSpec won't use syntax_suggest Abierto
Dificultad 4/5 3-5 días Aptitud para principiantes 35/100
ruby/syntax_suggest#171 · 3 comentarios ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 45/100
ruby/syntax_suggest#109 · 1 comentario ·
Todos los issues de ruby/syntax_suggest
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
TheOdinProject/curriculum#31417 · 2 comentarios ·
-
Allow faraday-http-cache 3.x Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
glossarist/glossarist-ruby#238 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
palladius/rails8-app-on-gcp#145 ·