Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Create RBS::Location lazily, like Prism (prototype: -14.5% retained memory)

オープン
#3,201 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
22/100
issue の種類
リファクタリング
明瞭さ
おおむね明確
活発さ
活発
技術スタック
c, ruby

調査の方向性

This is a design discussion, not a starter patch: read the prototype at github.com/ruby/rbs/compare/master...ksss:rbs:lazy-location-prototype, especially ext/rbs_extension/location_table.c, templates/ext/rbs_extension/ast_translation.c.erb, and lib/rbs/lazy_location.rb. Reproduce numbers with benchmark/memory_new_rails_env.rb and benchmark/memory_new_env.rb. Done means maintainers have answered the three questions (lazy locations vs Prism, Integer @location plus buffer:, map_type_name yield) and a follow-up implementation issue or PR is agreed—not implementing the full table-backed design from this thread alone.

索引モデルが issue の本文から書いたものです。

説明

Summary

RBS::Location objects take 22–28% of the memory retained by a loaded RBS::Environment, but almost none of them are ever read. I prototyped creating them lazily, following Prism's approach. In benchmark/memory_new_rails_env.rb, retained memory went down by 14.5% (-8.2 MB) and every location stays identical to master. This issue shares the numbers and the open design questions, to decide whether this direction is worth pursuing.

Motivation

The parser creates one RBS::Location per node, eagerly, and the node keeps it for its whole lifetime.

  • Locations are 2.38 MB of the 10.83 MB retained by benchmark/memory_new_env.rb (22%), and 15.76 MB of 56.66 MB in benchmark/memory_new_rails_env.rb (28%).
  • They are almost never read. With core + every stdlib library loaded, 60,734 Locations are created. I counted the ones whose contents were inspected (start_pos, [], to_s, ...):
    • after Environment.from_loader(loader).resolve_type_names: 0
    • after also building every instance/singleton/interface definition with DefinitionBuilder: 2 (both for an error message)

#location is called often, but mostly to pass the location to a rebuilt node (location: location in map_type_name, Environment#resolve_*, ...), not to look at it.

Prism does this already

The current parser follows Prism's design, and Prism (1.9) creates locations lazily:

# prism/lib/prism/node.rb
def location
  location = @location
  return location if location.is_a?(Location)
  @location = Location.new(source, location >> 32, location & 0xFFFFFFFF)
end

The C extension stores (start << 32) | length as an Integer, so no object is allocated until #location is called (except with freeze: true, where locations are created eagerly). Every node also keeps @source, and sub-locations like keyword_loc are separate node fields, each lazy in the same way.

What the prototype does

1. Nodes without child locations (Types::Bases::*, Optional, Union, Tuple, Record, Literal, Variable, Proc, Block, AST::Comment, AST::Annotation, Members::Public/Private, ...)

  • The C extension passes location: (start_char << 32) | length and buffer: buffer to .new.
  • #location builds the RBS::Location on first call and replaces the ivar.

2. Nodes with child locations (ClassInstance, Alias, Interface, Function::Param, MethodType, TypeParam, declarations, members, directives)

  • An Integer can hold only one range, so the child ranges go into a per-buffer table: a single int32 array attached to the Buffer (a C RBS::LocationTable). Each record is [node_type, start, end, (child_start, child_end) * N].
  • The node keeps the record index in @location and the buffer in @buffer.
  • Child names (:keyword, :name, ...) are fixed per node type, so they come from a static schema generated from config.yml and are not stored per location.
  • #location builds the same RBS::Location, with the same children, as the parser builds today.

3. Pass-through sites that rebuild a node of the same class (map_type, map_type_name, sub, update, Environment#resolve_*) pass the raw @location / @buffer along, so they don't materialize anything.

RBS::AST::Ruby::* (inline annotations) and the WASM path are left unchanged.

Results

ruby 4.0.6 / arm64-darwin, master 29b5984c. Both columns come from the same checkout; only the sources were switched and the extension rebuilt. Numbers are from memory_profiler.

benchmark/memory_new_rails_env.rb:

master prototype
Total retained 56.66 MB / 663,335 objects 48.45 MB / 503,125 objects -14.5% / -24.2%
Retained RBS::Location 161,671 1,178 -99.3%
Total allocated 90.67 MB / 1,232,424 objects 93.12 MB / 1,221,053 objects +2.7% / -0.9%

benchmark/memory_new_env.rb (core only):

master prototype
Total retained 10.83 MB / 99,116 objects 9.57 MB / 74,314 objects -11.7% / -25.0%
Retained RBS::Location 24,917 28

Where the retained memory goes in the Rails case:

RBS::Location objects removed -15.62 MB
RBS::LocationTable added +3.58 MB
Nodes that move to a larger GC slot because of the extra @buffer ivar +3.83 MB
↳ Members::MethodDefinition (8 → 9 ivars, 80 → 160 bytes) +1.62 MB
↳ Types::Function::Param (3 → 4 ivars, 40 → 80 bytes) +0.97 MB
↳ Types::ClassInstance (3 → 4 ivars, 40 → 80 bytes) +0.88 MB
↳ others (Types::Alias, AttrReader, Class::Super, ...) +0.36 MB
Net -8.21 MB

Step 1 alone (nodes without child locations) gives -3.53 MB (-6.2%) in the Rails case. None of those nodes cross a slot boundary, because they have at most two ivars before @buffer is added.

CPU time was not measured.

Correctness
  • Parsing every file in core and stdlib and dumping each node's location (range, plus the name, required/optional flag and range of every child) gives the same 60,280 lines with master and with the prototype.
  • rake test passes except RBS::WASM::SerializationTest (4 errors). That test compares the ivars of the C and WASM results directly, and the WASM deserializer was not updated.

Concerns and open questions

  1. @location can be an Integer. #location still returns an RBS::Location, but code that reads the ivar directly (instance_variable_get, Marshal, subclasses) sees an Integer.

  2. Constructors get a buffer: keyword, and code that rebuilds a node must pass @location / @buffer instead of location to keep the savings. Code outside rbs that does Foo.new(..., location: other.location) still works, but materializes the location.

  3. map_type_name yields the location as its second block argument (yield(name, location, self)), which materializes it for every type. Every block in lib/ ignores it (|name, _, _|). The prototype checks block.parameters and only materializes when that parameter is named. That keeps the API but allocates on every call, which is where the +2.7% allocated bytes comes from. Options: deprecate the argument, or yield something cheaper.

  4. Slot growth. Adding @buffer pushes some nodes into a larger GC slot (+3.83 MB above). Prism nodes always carry @source, @node_id and @flags, so the extra ivar costs little there; RBS nodes are smaller, so it shows. Possible ways out: keep the buffer reference somewhere other than an ivar for those classes, or reduce the ivar count of MethodDefinition.

  5. Table size. The table stores absolute int32 positions. Storing child ranges as uint16 offsets from the node's start would roughly halve its 3.58 MB.

  6. Location identity. Today a node and the copy made by map_type_name share one Location object. With lazy creation, each materializes its own, so a.location.equal?(b.location) no longer holds, and materializing both costs two objects.

  7. Frozen / Ractor-shareable ASTs. Lazy creation writes the ivar on first access, so it doesn't work with frozen nodes. Prism creates locations eagerly when freeze: true. RBS would need the same if it ever shares ASTs across Ractors.

  8. WASM / JRuby path. lib/rbs/wasm/deserializer.rb builds nodes in Ruby and would need the same representation (or keep eager locations).

  9. Signatures. sig/ needs @location: Location | Integer, the buffer: keyword, and so on.

Lifetimes don't change: a Location already keeps its Buffer alive, and now the node keeps it through @buffer. The table is attached to the Buffer, so it is freed together with it.

Possible plan

  1. Nodes without child locations only. Small API surface, no slot growth, -6.2% retained memory in the Rails benchmark.
  2. Table-backed locations for nodes with child locations. Around -14.5% in total, with concerns 3 and 4 to settle.
  3. Reduce the costs: slot growth, table compaction.

Questions:

  • Is aligning RBS::Location with Prism's lazy locations the direction you want?
  • Is it acceptable for @location to hold an Integer, and to add a buffer: keyword to node constructors?
  • For map_type_name, would you rather deprecate the location block argument, or keep it as is?

The prototype is three WIP commits: https://github.com/ruby/rbs/compare/master...ksss:rbs:lazy-location-prototype
The C side is ext/rbs_extension/location_table.c and the template changes in templates/ext/rbs_extension/ast_translation.c.erb; the Ruby side is lib/rbs/lazy_location.rb plus mechanical changes to the node classes.

Related: #3200 (sharing the frozen empty array for empty node lists), another memory reduction for retained ASTs.

主要言語
Ruby
スター
2.2k
フォーク
258
平均マージ
1日 8時間
マージ済み PR(30日)
44

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

ruby/rbs のほかの issue

ruby/rbs の issue をすべて見る

似ている issue

Ruby の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。