Originally created by: pyy3
Type: Spec / design (Phase 3 — Scale Architecture §2; last remaining Scale item)
Implementation repos: konsolidat (dbt / ClickHouse) + konsol (profiles/orchestration)
The pipeline is validated on a single-node ClickHouse. At 50–500 legal entities a single node can't hold GL history with acceptable latency. The canonical grain already carries entity_id + erp_source but nothing shards on it. (Scale §1 incremental extraction shipped in konsolidat#44; §3 per-connector health shipped in konsol#25/#27 — sharding is what's left.)
entity_id across a cluster using Distributed + ReplicatedMergeTree; shard key cityHash64(entity_id) so a given LE colocates._local ReplicatedMergeTree per shard; the unsuffixed name becomes the Distributed table the API + Cube.js query. Partition keys preserved per shard (sharding ⊥ partitioning).ON CLUSTER DDL.cluster: konsol_cluster in profiles.yml; cluster mode behind a profile/var so single-node stays the default deploy.gold_trial_balance) for correctness on Distributed reads; all 144 dbt tests green in cluster mode.cityHash64(entity_id) — do we need a composite key (entity_id + fiscal_year) for the largest tenants?Not verifiable on the current single-node stack — needs a multi-node ClickHouse + Keeper. This issue is build + document + cluster-mode test; live end-to-end verification requires a real cluster.
Prior design detail: docs/prd/PRD-SCALE-ARCHITECTURE.md §2.
Ticket changed by: grynn-in