Decision: where tag associations are stored (Phase 3 tag-based masking)
Status: DECIDED (2026-06-19). Scopes the storage layer for tag-based masking (the Snowflake tag-masking parity pillar). Pairs with ranger-fine-grained-service-type.md and fine-grained-policy.md.
The two halves of a tag system
Tags split into two independently-stored things. Conflating them is the usual mistake.
- The mask-per-tag RULE (“any column tagged
PII-> mask show-last-4”). - The tag-to-column ASSOCIATION (“column
sales.customers.ssnhas tagPII”).
Decision
-
Rule (1) lives in Apache Ranger. Ranger’s
tagservice holds tag-based mask / row-filter policies, and SQE’sRangerStoredownload bundle already returnstagPolicies. No change to the policy-source model: the same rules are shared with Spark/Kyuubi, exactly like our resource policies. -
Association (2) lives in Iceberg/Polaris table metadata as a single namespaced table property
sqe.column-tags, a JSON object mapping column name to a list of tags:sqe.column-tags = {"ssn": ["PII"], "amount": ["FINANCIAL"]}Stored in
TableMetadata.properties(an arbitrary key/value map that Polaris persists and SQE already reads viatable.metadata()). The user-facing surface is nowALTER TABLE ... SET TAGS / UNSET TAGS(with the SnowflakeMODIFY|ALTER COLUMN ... SET TAGforms), not rawSET TBLPROPERTIES; the DDL writes this one property for you. -
Cross-engine (Spark) tag enforcement is an OPTIONAL one-way sync, not a dependency: a Phase-3.1 job mirrors
sqe.column-tagsinto Ranger’s tag store (/service/tags/...) so Kyuubi’s Ranger plugin honors the same column tags. SQE never depends on this sync being run.
Why Iceberg/Polaris properties for the association, not the Ranger tag store
Four deciding factors, in order:
-
Federated catalogs. Per ranger-fine-grained-service-type.md, SQE’s engine-side enforcement is the ONLY fine-grained layer that covers external / federated catalogs (Polaris cannot gate those). Tags-as-table-properties work for ANY Iceberg table SQE can read, federated included. Populating Ranger’s tag store per federated resource (with exact name matching) is fragile and would leave federated tables untagged.
-
No Atlas, no tagsync. Ranger tag associations normally come from Apache Atlas via tagsync. There is no Atlas for Polaris/Iceberg here, so a Ranger-only association store would sit empty unless something pushes to it. Iceberg properties need no Atlas and no tagsync.
-
Tags travel with the data. Properties live in the table metadata, so they survive clone, replicate, and rename. Ranger associations are keyed by resource name and break on rename/move.
-
SQE reads it natively. SQE already loads
table.metadata()on every scan; reading one extra property is trivial. No new client, no new store.
The mask-per-tag RULE (the genuinely valuable shared-with-Spark part) still lives in Ranger, so the cross-engine policy story is preserved. Only the association source of truth moves to the data.
Tradeoff accepted
Spark/Kyuubi read tag ASSOCIATIONS from Ranger’s tag store, not from Iceberg
properties. So out of the box, Spark will not honor sqe.column-tags-sourced
tags. That is what the optional Iceberg -> Ranger sync (above) is for, and it is
only needed when Spark must enforce the same column tags. For SQE-native
governance (the common case, and the only option for federated catalogs), no
sync is required.
Storage format detail
- One property
sqe.column-tags(a single JSON blob), not one property per column. Atomic to read/write, no per-column-property support needed (Iceberg per-column metadata beyonddocis not in the spec). Columndocis left for human descriptions. - Tag names are opaque strings that must match the resource/tag names used in
Ranger
tagPolicies. - iceberg-rust 0.8.0:
TableMetadata::properties() -> &HashMap<String,String>for the read; the write goes through the catalog’s update-table-properties path (the same mechanism CTAS/ALTER use). Confirm the exact iceberg-rust write API at implementation time.
Resolution path at query time (Phase 3 build)
- On scan, read
sqe.column-tagsfrom the table metadata -> map column -> tags. - From the
RangerStorebundletagPolicies, resolve the mask/row-filter for each tag that applies to the user’s roles (same matching as resource policies). - Feed the resulting masks/filters into the existing
PolicyEnforcer/PlanRewriter. Resource policies still win on conflict per Ranger ordering.
This reuses the entire enforcement path already shipped in Phase 1/2A; Phase 3 is the tag SOURCE + the tag-policy resolution, not new enforcement.
Operating note: tag row filters need a Ranger flag (validated 2026-07-31)
Both halves above work against a live Apache Ranger 2.8, proven end to end by
crates/sqe-coordinator/tests/it/access_control_e2e.rs (make test-access-control). Two behaviours are worth knowing before you deploy this.
Tag mask types are component-qualified. Ranger’s tag service definition
does not define bare mask names. It aggregates the mask types of every component
that can be decorated, so the entries are hive:MASK_SHOW_LAST_4,
hive:CUSTOM, trino:MASK_NULL and so on. SQE reads a hive-type service, so
it accepts the bare form and the hive: form; a mask authored under another
component’s prefix is deliberately left unmatched, which restricts the tagged
column (fail-closed) rather than applying another engine’s policy.
Tag ROW FILTERS are off by default, and it is not a version limitation.
Ranger propagates each component’s dataMaskDef into the tag service definition
unconditionally, but propagates rowFilterDef only when Ranger Admin runs with
<property>
<name>ranger.servicedef.autopropagate.rowfilterdef.to.tag</name>
<value>true</value>
</property>
in ranger-admin-site.xml (AbstractServiceStore, default false). Without it
the tag service definition carries a populated dataMaskDef and an empty
rowFilterDef: {}, and Ranger rejects a tag row-filter policy with:
tag policy can specify values for one of the following resource sets:
does not have any resource hierarchies
Set the property, restart Ranger Admin, then re-save a component service
definition so the propagation runs. Tag row filters then behave exactly like
resource row filters in SQE: resolve_tag_policies returns them keyed by tag and
the rewriter ANDs them above the scan.
The e2e suite cannot assume an operator-configured Ranger, so its fixture patches
the capability in over the REST API instead
(ranger_fixture::ensure_tag_rowfilter_support). That is a test-environment
shortcut, reset by a Ranger upgrade or a volume wipe. Production deployments
should use the property.
One trap if you ever PUT the tag service definition yourself: Ranger’s own
aggregate does not round-trip through Ranger’s validator. On 2.8.0 with the stock
component set it carries a duplicate ozone:assume_role access type (duplicate
itemId 201209) and elasticsearch implied grants naming access types the
definition never declares, so a verbatim re-submit is rejected. Deduplicate and
prune those first.