kudu维表join 出来的数据会有重复

Open
#346 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
25/100
Issue type
Bug
Clarity
Needs clarification
Activity status
Stale
Tech stack
sql

Research direction

Start with the provided Kudu dimension-table DDL and reproduce the join against d_goods_wlm_sku_rel. Inspect the resulting rows and determine whether duplicates come from the join semantics, source keys, or caching configuration. Done means the expected join result contains no unintended duplicate records, but the issue needs a query and expected output before work can begin.

Written by the indexing model from the issue text.

Description

---kudu维表语句
CREATE TABLE side_rt_d_spu_sku_relation(
spu_id int,
sku_ids varchar,
PRIMARY KEY(spu_id),
PERIOD FOR SYSTEM_TIME
)WITH(
type ='kudu',
master ='xxxxxx'
tableName='d_goods_wlm_sku_rel',
cache ='LRU',
cacheSize ='10000',
cacheTTLMs ='60000',
parallelism ='1',
partitionedJoin='false'
);

Dominant language
Java
Stars
2k
Forks
913
PR merge metrics
No merged PRs in 30d

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from DTStack/flinkStreamSQL

All issues in DTStack/flinkStreamSQL

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.