Menu ▾ ▴

#57 Spec: make ClickHouse cluster mode functional (rework after #56 scaffolding)

open
nobody
None
2026-09-15
2026-06-16
Anonymous
No

Originally created by: pyy3

Type: Spec / follow-up (Phase 3 — Scale §2)
Implementation repo: konsolidat (dbt + ClickHouse cluster config)
Requires: a live multi-node ClickHouse cluster to validate — none of this is verifiable single-node.

Context

[#56] (merged, closed [#53]) landed cluster-sharding scaffolding, gated behind cluster_enabled (default off). The single-node default is provably byte-identical and safe. But cluster mode is not yet functional — turning cluster_enabled on today will not produce a working sharded cluster. This issue tracks the rework to make the ON state real.

Remaining work

  • [ ] Materialization rework (the core). Model-level config(cluster=...) is inert under dbt-clickhouse 1.10 — ON CLUSTER is driven by the profile cluster: key, and plain incremental/table materializations never build a <model>_local table, so the hand-rolled create_distributed_tables run-operation has no _local to overlay. Switch the cluster-aware models to the adapter's native distributed_table / distributed_incremental materializations (which manage _local + Distributed + sharding_key from model config), and retire or repurpose the run-operation. Shard key is cityHash64(data_area_id).
  • [ ] ClickHouse Keeper per-node config. clickhouse/cluster/keeper.xml hardcodes server_id=1 while docker-compose.cluster.yml mounts the same file on all nodes → broken quorum. Each node needs a unique server_id; each shard needs its own macros.xml (macros-shard{N}-replica{M}.xml). Add a generator or per-node templating; only shard1 macros exist today.
  • [ ] Stand up + validate on a real cluster (≥2 shards + 3 Keeper nodes; docker-compose.cluster.yml is the scaffolding). Acceptance criteria (from PRD §2):
  • SELECT count(distinct data_area_id) FROM epm_gold.gold_trial_balance equals the single-node baseline (no rows lost / double-counted across shards).
  • Rows for one data_area_id reside on exactly one shard (shardNum() check).
  • All 144 dbt tests green in cluster mode.
  • Cross-shard gold aggregations correct: consolidation, IC eliminations, NCI, FX translation.
  • [ ] Decide the open questions (PRD §2): hot-shard key for a very large single LE (composite data_area_id + fiscal_year?); reshard-vs-skew policy on new-LE onboarding; whether Cube.js pre-aggregation needs a shard-aware refresh key.

Done criterion

--vars '{"cluster_enabled": true}' + a cluster: profile target produces a correct, tested sharded warehouse, and the single-node default remains byte-identical when off.

Prior scaffolding + design detail: PR [#56], dbt_project/macros/cluster.sql, docs/admin-guide/cluster-setup.md, docs/prd/PRD-SCALE-ARCHITECTURE.md §2.

Related

Tickets: #100
Tickets: #53
Tickets: #56

Discussion

  • Anonymous

    Anonymous - 2026-07-01

    Originally posted by: grynn-in

    Decision (next round, P3 / scale-gated). Scaffolding exists (#56/#100). Making cluster mode functional only pays off at real scale/HA needs, which we don't have yet on a single-node dev/demo. Defer until a concrete scale or availability requirement lands; revisit alongside [#90] (prod infra). Keeping open.

     

    Related

    Tickets: #90

  • Anonymous

    Anonymous - 2026-07-01

    Originally posted by: grynn-in

    📋 Decision brief (options + trade-offs + recommendation): docs/developer-guide/decisions/konsolidat-57-clickhouse-cluster.md — merged in [#133].

     

    Related

    Tickets: #133

  • Anonymous

    Anonymous - 2026-09-15

    Originally posted by: grynn-in

    Verdict: LIVE as a spec, unbuilt — but worth re-confirming against the multi-tenant direction before anyone starts it.

    Still accurate: cluster_enabled remains in dbt_project.yml (default false) with the cluster macro file in place, so this is still scaffolding awaiting the real work, and still unverifiable without a multi-node cluster.

    The thing that has changed is the target shape. With the dbt project moving into the Frappe app and multi-tenancy arriving through Frappe's multi-site model rather than through warehouse sharding, cluster mode may be solving a scale problem the new architecture addresses differently. Not stale — but it should be re-confirmed as wanted before it is scheduled.

    🤖 Triage against main — Claude Code · https://claude.ai/code/session_01P3Pf9835FeLeXjRrYTTZ1M

     
  • Anonymous

    Anonymous - 2026-09-15

    Originally posted by: grynn-in

    Re-scoped — re-confirm the target before scheduling

    Still accurate as a spec: cluster_enabled remains in dbt_project.yml (default false) with the cluster macro in place, so this is scaffolding awaiting the real work, still unverifiable without a multi-node cluster.

    What has changed is the target architecture. The direction of 15 Sep 2026 is that the dbt project moves into the Frappe app and multi-tenancy arrives through Frappe's multi-site model, with the ELT/ETL layer left out. Cluster sharding may be addressing a scale problem that shape solves differently — or it may still be needed for a single very large tenant, which is a different justification from the one in the body.

    Recommendation: treat this as blocked on a decision rather than on effort. Confirm whether multi-node ClickHouse is still on the roadmap under the new architecture, and record the answer here. If yes, the spec stands as written; if no, it closes.

    🤖 Re-scoped 15 Sep 2026 against main — Claude Code · https://claude.ai/code/session_01P3Pf9835FeLeXjRrYTTZ1M

     

Log in to post a comment.