--json output is absurdly huge
還沒有人認領這個 Issue。
評估
研究方向
首先檢視 --json 輸出如何表示 changedOperations,以及相應的 HTML、Markdown 和 stdout 格式。比較 issue 中描述的重複物件、深度巢狀的已變更屬性和省略的詳細資訊,然後確定統一輸出模型的範圍。完成的標準是形成一項取得共識的設計,在減少重複的同時保留跨格式的完整 diff 資訊。
由索引模型根據 Issue 內容生成。
描述
I wrote this on #215 day before yesterday:
... the [json] file size was over 10,000x that of the html and md files. No joke: the same diff produced a 113KB html file and a 1.1GB json file 🤯 I've had some difficulty doing anything with it to see what data it contains...
The "difficulty" was that all tools I had on hand for analyzing json were getting OOMKilled before they could finish parsing the file. They all ate up a full 9.8GB of RAM within about 45 seconds and died after a few minutes.
Luckily I learned about the --stream option for jq. It still ate up every free byte of memory that it could and took about five minutes to do anything, but it didn't crash and it got me started.
This SO answer helped me figure out what fields to look up using a generalization of this other answer. Between the two of these, I managed to extract a single endpoint of interest from changedOperations into a file of 729KB--which is almost 2x as large as the entire original spec.
I specifically was trying to find why the html diff was telling me an endpoint had a breaking change but showed identical schema. Turned out one field's maxLength property was reduced, but the html (and md) output doesn't include that level of detail. But more importantly: while investigating that, I noticed a number of objects and structures were being repeated in various places.
One in particular jumped out at me, so I wrote a little more jq to count up the number and location of occurrences:
[ path(.. | select(type == "object" and has("context"))) as $p
| { "path": ( $p | join(".") ), "context": ( getpath($p + ["context"]) ) }
]
| group_by(.context)
| [.[0][].path] as $c0paths
| [.[1][].path] as $c1paths
| [.[2][].path] as $c2paths
| [ {"context": (.[0] | first | .context)
, "no_paths": ($c0paths | length)
, "paths": ($c0paths)}
, {"context": (.[1] | first | .context)
, "no_paths": ($c1paths | length)
, "paths": ($c1paths)}
, {"context": (.[2] | first | .context)
, "no_paths": ($c2paths | length)
, "paths": ($c2paths)
}
]
[
{
"context": {
"url": "/authenticate",
"parameters": {},
"method": "POST",
"response": false,
"request": true,
"required": true
},
"no_paths": 640
},
{
"context": {
"url": "/authenticate",
"parameters": {},
"method": "POST",
"response": true,
"request": false
},
"no_paths": 14
},
{
"context": {
"url": "/authenticate",
"parameters": {},
"method": "POST",
"response": false,
"request": true
},
"no_paths": 14
}
]
(paths arrays ommitted for brevity.)
All told, that accounts for something like 74KB of duplicated data.
I also like this one:
{
"compatible": false,
"incompatible": true,
"unchanged": false,
"different": true
}
These four fields appear in well over 1,000 places and account for around 82KB. I see an easy way to cut that down to around 41KB...
The operation objects from the original specs are included in their entirety under two top level keys. Then, they're each duplicated under two other keys, accounting for 12KB altogether. Since someone making a diff would presumably have both specs on hand, it's unnecessary to include these full models at all, nevermind four times (that I've found so far).
There are a number of other structures that pop up all over the place, hundreds of times. I don't know how many are exact duplicates of each other, but I'm guessing it's a lot more than is strictly necessary.
I suspect the cause to be that Java objects may be getting serialized into JSON verbatim, including their references to related objects. I'm sure those references are necessary in the Java objects themselves and that they make the library very easy to work with. However, JSON output isn't for a library consumer: the Java object model doesn't make sense in other contexts.
The json output could just use references itself, but, honestly, please don't. Json refs are about as hard to parse as a 1.1GB file, so getting a 1,000-fold disk savings would be just about offset by the need to get tools that can handle refs correctly and also deal with large file sizes.
Instead, I would suggest a better model for the output.
- Since this is a diff tool, there's no reason to include entire models from the original specs. We only need the things that are actually different.
- As little duplication as possible. For example, there is no conceivable reason to include that
contextobject over 600 times when there were only three unique contexts--and the three of them share half of their fields. It makes sense in Java; it makes no sense in a file that's to be used by humans (give or take some automated parsing). - Put interesting data closer to the top of the object tree. Here's one
jqpath (of several) to get at the changedmaxLengthproperty:.changedOperations[0].requestBody.changedElements[1].changedElements[0].changedElements[0].maxLength. This is almost impossible to review by eyeball (even with pretty-print) and it's a pain to parse out even with tools likejqor nushell's tables and dataframes.
In addition, I would suggest that such a new model be unified across all output formats. Why should html have less information than md, which has less information than stdout, which has less information than json? I just want to get back to writing my API client code knowing that the diff I'm using for guidance is comprehensive and correct, not worry if I missed an obscure property like maxLegth because it was omitted from the output, nor spend a day learning an FP-paradigm DSL just to figure out that's what was omitted!
And as a last thought, which probably deserves its own issue: --yaml option. Compliant YAML parsers should be stream-oriented and capable of parsing in chunks, meaning the issue of running out of memory can be side-stepped even for files that are several GB in size-- unlike json, which must be parsed in full to be understood by the application.
- 主要語言
- Java
- 星號
- 1.1k
- 分支
- 190
- PR 合併指標
- 30 天內沒有已合併 PR
環境準備
- 提供 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
OpenAPITools/openapi-diff 的其他 Issue
-
enhancement
難度 2/5 1-3 小時 新手友好度 68/100
OpenAPITools/openapi-diff#506 ·
-
good first issue help wanted
難度 2/5 1-3 小時 新手友好度 68/100
OpenAPITools/openapi-diff#364 ·
-
bug OpenAP 3.1.0 Support
難度 3/5 1-2 天 新手友好度 68/100
OpenAPITools/openapi-diff#910 · 1 則留言 ·
-
Render capabilities
難度 3/5 1-2 天 新手友好度 55/100
OpenAPITools/openapi-diff#893 · 1 則留言 ·
-
[Bug] Backward compatibility check fails on reordered discriminator mappings可能已有人在做 @MoChiUaena 於 8 天前認領。 未關閉Breaking/Non-Breaking classification
難度 3/5 1-2 天 新手友好度 55/100
OpenAPITools/openapi-diff#886 ·
查看 OpenAPITools/openapi-diff 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 72/100
-
area/docs
難度 1/5 1 小時以內 新手友好度 88/100
維護者通常 1 天內回覆
-
BoxChart rejects valid List.of data with NullPointerException可能已有人在做 @PHJ2000 今天認領。 未關閉
難度 2/5 1-3 小時 新手友好度 76/100