Performance idea: Partition before executing current uniform indentation search
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 20/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- ruby
- Lĩnh vực
- performance, tooling
Hướng nghiên cứu
Bắt đầu với test đang thất bại trên branch schneems/partition và lần theo các lời gọi Ripper.parse của thuật toán tìm kiếm hiện tại. Đọc cách thụt lề và các cặp kw/end được xử lý, sau đó so sánh việc phân vùng hoặc các bước mở rộng lớn hơn với hành vi hiện có. Hoàn thành khi trường hợp chín nghìn dòng tránh được timeout mà vẫn duy trì chất lượng của kết quả lỗi cú pháp.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
This is a failing test: https://github.com/zombocom/dead_end/tree/schneems/partition. The file is nine thousand lines and it takes a tad over 1 second to parse which means it hits the timeout.
We are already fairly well optimized for the current algorithm so to be able to handle arbitrarily large-sized files we will need a different strategy.
The current algorithm takes relatively small steps in the interest in producing a good end result. That takes a long time.
Here's my general idea: We can split up the file into multiple large chunks before running the current fine-grained algorithm. At a high level: split up the file into 2 parts and see which holds the syntax error. If we can isolate the problem to only half the file then we've dropped processing time in half (relatively). We can run this partition step a few times.
The catch is that some files (such as the one in the failing test cannot be split without introducing a syntax error (since it starts with a class declaration and ends with an end). To account for this we will need to split in a way that's lexically aware.
For example on that file, I think the algorithm would determine that it can't do much with indentation 0 so it would have to go to the next indentation, there it could see there are N chunks of kw/end pairs, it could divide into N/2 and see if one of those sections holds all of the syntax errors. We could perform this division several times to arrive at a subset of the larger problem, then run the original search array on it.
The challenge is, that we will essentially need to build an inverse of the existing algorithm. Instead of starting with a single line and expanding towards indentation zero, we'll start with all the lines and reduce towards indentation max.
The expensive part is checking code is valid via Ripper.parse, sub dividing large files into smaller files can help us isolate problems sections with fewer parse calls, but we've got to make sure the results are as good.
An alternative idea would be to use the existing search/expansion logic to perform more expansions until a set of N blocks are generated then check all of them at once. Then once the document problem is isolated, go back and re-parse only the N blocks with the existing. Algorithm. (Basically the same idea as partitioning, but we're working from the same direction as the current algorithm, just taking larger steps (which means fewer Ripper.parse) calls. However we would still need a way to sub-divide the blocks with this process in the terminal case that the syntax error is on indentation zero and the document is massive and all within one kw/end pair.
- Ngôn ngữ chính
- Ruby
- Star
- 350
- Fork
- 17
- Merge trung bình
- 48 phút
- Pull request đã merge (30 ngày)
- 5
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của ruby/syntax_suggest
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 38/100
ruby/syntax_suggest#258 · 6 bình luận ·
-
Accidental if instead of a blockĐang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 30/100
ruby/syntax_suggest#206 ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 42/100
ruby/syntax_suggest#205 · 1 bình luận ·
-
RSpec won't use syntax_suggestĐang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
ruby/syntax_suggest#171 · 3 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 45/100
ruby/syntax_suggest#109 · 1 bình luận ·
Tất cả issue của ruby/syntax_suggest
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
simp/pupmod-simp-ssh#246 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
simp/pupmod-simp-pupmod#261 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
-
documentation_basics detective marks OSPS-DO-01.01 Unmet for projects whose user guide is the READMEĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
ossf/best-practices-badge#3054 ·
Maintainer thường phản hồi trong vòng 1 ngày