The open-source semantic layer that infers itself — skip the quarter of hand-writing dbt YAML.

Point it at your warehouse. It profiles every column, discovers the foreign keys nobody declared, decodes the status columns, finds the business rules hiding in your aggregate tables, and writes the whole thing down as an open, portable semantic layer — with confidence, provenance, and lifecycle on every single claim — ready for any AI agent to consume over MCP.

pip install -e ".[warehouses]"

semlayer init snowflake # generates the minimal-grant setup script

semlayer infer snowflake -o layer.yaml \

--context ./docs/ --context ./etl-repo/CLAUDE.md # optional: your wikis/dictionaries as priors

semlayer review layer.yaml # accept/reject what the engine inferred

semlayer mcp layer.yaml # serve it to Claude, Cursor, or any MCP client

semlayer lint layer.yaml query.sql # check any SQL (yours or an agent's) against the layer

semlayer drift layer.yaml snowflake # catch schema changes (cron- and CI-friendly)AI agents fail on real warehouses: frontier models solved just 21.3% of Spider 2.0's enterprise-warehouse tasks at publication (vs ~91% on the earlier academic Spider 1.0) — and even today's best agentic scaffolds only reach ~30%. The fix is a semantic layer — but every existing tool (dbt, LookML, Cube, Snowflake semantic views) makes humans write it by hand, and it goes stale the day someone runs an ALTER TABLE.

On our messy-warehouse benchmark (cryptic names, zero declared constraints, hidden business rules), an agent using the inferred layer answers 87% of business questions correctly vs 42% from the raw schema (+107% relative), 89% when it also runs the layer's SQL linter — and the errors it fixes are the silent kind: raw-schema "total revenue" happily sums cancelled orders; a fan-out join quietly triples a total. Full benchmark, methodology changes included, and where we DON'T help →

Everything lands with confidence (calibrated, with the misses published), provenance (which signals produced it), and lifecycle (inferred → reviewed → certified, plus deprecated/orphaned), governed by a normative consumer contract that makes silent misuse — summing across a fan-out, joining SCD2 at current-row, filtering on guessed decodes — non-conforming, not merely unwise.

- ~$0.70 per 100 tables end-to-end on the cheap model tier, with your own API key (measured cost model). LLM calls are escalation-only; ~80% of columns resolve from statistics alone.

- --no-llm: fully deterministic mode, zero API calls, still useful (0.780 typing / 0.834 role across 7 fixtures, against 0.809 / 0.838 with the model) — for orgs where LLM access needs procurement.

- --no-sample-egress: no value read from a warehouse cell is sent to the LLM, at no net accuracy cost across our 7 fixtures (0.809 typing / 0.842 role, versus 0.809 / 0.838 with samples). Per warehouse it trades rather than being uniformly free — up to a point better or worse. Re-measured 2026-10-03; see docs/cost-model.md.

- Read-only, minimal-grant: semlayer initgenerates the grant script; the live test suite (tests/test_snowflake_live.py, runnable against your own account) proves the reader persona cannot write.

- Telemetry: anonymous command counts spooled locally only — nothing leaves your machine in this release; opt out with SEMLAYER_TELEMETRY=off. (details)

- Best on messy warehouses. On clean, well-named schemas our benchmark shows raw DDL is already sufficient — we publish that negative result rather than hide it. If your warehouse is tidy TPC-DS, you may not need us.

- Warehouses: Snowflake, BigQuery, DuckDB (+ Iceberg on S3 via the DuckDB bridge). Exporter: dbt (losses reported, never silent). LLM: Anthropic API (Bedrock/Vertex routing is next).

- Not included: hosted service, ontology enrichment (it failed our own ablation gate — receipts), LookML/RDF exporters.

- The eval harness ships in this repo — fixtures, gold layers, competency questions, benchmark runner. Run our numbers yourself: python fixtures/build.py && pytest tests/ -q.

Running the beta? File the Beta feedback form (attach the *.report.json the CLI writes next to your output — timings and counts only, never your schema or data), or email hello@semlayer.dev for anything sensitive.

spec/ format schema + consumer contract · src/semlayer/ the engine · fixtures/ 9 eval warehouses + golds + CQ suites · docs/ benchmark, cost model, spike reports · ARCHITECTURE.md how it works

License: Apache-2.0.