Skip to main content

Market data — the semantic API

What changed, and why

Driftless used to reach GTM Fábrica over a private HTTP Query API that served two tools to the model: describe_schema and query_sql_readonly. The model wrote SQL. That interface required a model to hold the warehouse’s entire semantics in its head on every call — schema, joins, cardinalities, JSONB shapes, geographic normalization, strong versus approximate identity, currency, amount granularity, temporality, coverage, licensing, and the limit on which contacts may be exhibited. The observed cost was wrong SQL, unnecessarily expensive searches, timeouts, blended currencies and granularities, weak joins treated as identity, lexical false positives, misread procurement semantics, calls burned repairing queries, and a skill that grew to compensate.
SQL is an interface for analysts. The semantic API is the interface for agents.
The model now chooses a typed business operation. It does not choose SQL, joins, scopes or any critical semantics. The current domain protocols, targeted Web Evidence boundary, Workbench and evaluation/deletion gates are consolidated in commercial-intelligence-v1.md.

The route

There is no HTTP between the two repos, no MCP hop, no gRPC, no second microservice, no second pool, and no path from a model to SQL. The agent tools and the HTTP controllers call the same services in process — the tools do not call the controllers over HTTP. Every read runs inside one transaction:
REPEATABLE READ pins one snapshot so rows and the coverage they are reported against cannot come from different corpora. READ ONLY is a second, independent guarantee alongside the database role. The timeout is LOCAL, so one slow operation cannot widen the budget for the next user of that pooled connection.

Ownership

Endpoints

All workspace-scoped. None is @Public().
GET award-suppliers/:rfc/history is retired. History moved to POST awards/history alongside every other multi-row operation, because a history that pages needs the same cursor/limit body every search already takes — a path parameter had nowhere to carry either. The route is gone, not redirected: a caller still on the old path gets a 404, not a compatibility shim. GET capabilities accepts an optional facets query parameter (comma-separated dimension names, e.g. ?facets=procedure_type,mark_kind) that restricts both the computation and the response to the named corpus dimensions; an unrecognized name is refused as unknown_facet with the ones that exist. GET capabilities/compact is the same bounded, source-aware projection the agent runtimes are injected, for an integrator that wants it without the full observed-value catalogue.

Agent tools

The market surface contains thirteen typed tools: free discovery through market_capabilities, plus twelve metered warehouse operations: search_suppliers, get_supplier, count_suppliers, compare_segments, search_opportunities, get_opportunity, search_awards, get_supplier_history, aggregate_awards, search_risks, screen_risks, and search_permits. search_risks answers “is this party carrying a published adverse mark?” — an EFOS barred-supplier listing or a sanción/inhabilitación (mark_kind: efos | sancion), never a judicial conviction and never the whole of a party’s regulatory history: a party can carry several marks, including published exonerations. rfc is the only strong identity in this relation and is structurally validated before it reaches the corpus; entity_name is always approximate and is resolved EXACT-ONLY — a name with no exact published spelling (once accents and case are folded) is refused with entity_name_not_exact and the similar names as candidates, never searched by similarity on the caller’s behalf. It also has no free-text search document behind it the way opportunities and awards do. search_permits answers “does this party hold a published permit, concession or capex commitment?” — a granted right recorded at load time, never a live operational verification that a plant is generating or a concession is being worked today. Identity runs the other way from every other relation here: holder_rfc is strong where published, but the two live sources (energia-cne, concesiones-mineras) publish no RFC on almost any row, so holder_name — resolved to an exact published spelling, and refused with candidates when none matches, exactly like supplier_name — is the join this relation was actually built around. aggregate_awards accepts order_by: total_amount | award_count (default: total_amount). A group keyed by supplier_rfc also returns supplierNameLabel, the most frequent published spelling inside that RFC group. The label is for display only: grouping, counting and identity remain RFC-based, and rows without an RFC are excluded from an RFC ranking. Coverage is not a separate tool: it rides on every response. market_capabilities is the free discovery tool and is also injected by the runtime where the platform already has it. Capabilities are now TWO layers. The semantic contract — operations, identity, amounts, dates, contacts, matching — is code-owned and versioned with the schema. The value lists beside each exact filter are read from the LIVE licensed corpus on every load, in one bounded GROUPING SETS query per physical relation, and stamped with the corpus basis they were read against. The 6.1M-row directory is never scanned for facets: its dimensions are declared not_enumerated, because the frozen profile could not complete an exact field profile over it and a sample presented as a census would be worse than no list. The HTTP route returns the complete lists; the runtime is injected a compact, source-aware rendering that declares anything it dropped. Making a model spend a call to discover a system the platform already knows is ceremony, and on a bounded budget ceremony is the difference between an answer and a partial. The model never receives: a SQL schema, table names, describe_schema, query_sql_readonly, the ability to submit SQL, join rules, private fields, a DSN, or a database credential.

Separate research capability: search_web_evidence

The thirteen market tools above read the governed warehouse contract. search_web_evidence is not a fourteenth market tool; it is a separate research capability for what the warehouse structurally cannot hold — recency, an announced expansion, a private project, a corporate publication, or external confirmation that something is still active. It is a capability, never a provider. The domain never names an executor, and nothing about one may reach a prompt, a ResearchReport, a citation label or the frontend. Two mechanical guards enforce it: research-providers/research-providers.architecture.spec.ts and cognitive/market-research/market-research.architecture.spec.ts.
There is no provider configuration, no callback URL, no model-authored output schema, no discovery depth and no match limit — those fields do not exist on the type, so they cannot be passed. A request with no entity and no claim, or an objective too general to target a source, is refused as web_query_too_broad with a concrete correction before anything is spent. The result is provider-neutral and citable:
search_id, evidence_id, observed_at and source_domain are platform-owned, exactly like coverage and provenance on a warehouse envelope. A URL a model wrote is never accepted as provenance: the model cites an evidence id and the platform resolves the record from its own ledger. A claim’s status describes retrieval, not truth. supported means at least one citable source came back and none of them denied the claim; contradicted is absorbing and can never be overturned by agreeing sources; unverified means nothing citable came back. Contradiction detection is a closed, declared marker list, not a model call — a probabilistic answer to “may this be reported as support” is worse than no answer. Budget: two searches per run, on a pool market-data calls never consume. The third is refused with web_budget_exhausted and still recorded in the audit — a refused attempt that leaves no row reads as a model that simply stopped asking. Every attempt is audited under the operation name search_web_evidence, in the same queryAudit as warehouse calls, so a run is reconstructible from one place. Web failures are typed — web_search_unavailable, web_query_too_broad, web_budget_exhausted, web_results_unverified, web_provider_timeout — and none of them is critical. A web gap that stays open is a gap; it never removes a warehouse row and never turns an answered question into “no answer”. A deployment with no executor configured is a supported deployment: the tool is not offered, the web method is not loaded into the skill, and the run says which gap it could not close. Exhaustive discovery and contact enrichment are unreachable from this capability. It holds a port whose whole surface is capabilities/estimate/health/search/fetch/extract/cancel and calls exactly one of those; a recording proxy in research-providers/web-evidence.service.spec.ts asserts which methods are touched, so the guarantee is a runtime fact and not only a source scan.

The envelope

coverage, corpus_basis, semantic_warnings and provenance are platform-owned: readable by the model, writable only by the layer. Those are precisely the fields a model would shade if it could reach them, and coverage in particular is the only thing that separates “this query returned nothing” from “this market does not exist”. Every monetary value — a row’s amount, a history or aggregate total, a min/max, a period total or delta, a risk mark’s fine — is a decimal STRING at two places, computed and rounded in numeric and serialized with ::text. No amount is ever parsed into a JS number on the way out, and a caller should not parse one on the way in beyond the precision it needs: two decimals is the money the publisher wrote, and the digits past them that the warehouse carries are a loader artifact, not a published fact. Counts, ranks, scores, percentages and physical magnitudes (capacity_mw, investment_mdd, surface_hectares) are not currency amounts and are served exactly as published.

Agent Observation — one truth, three derivations

The envelope is the right shape for an API client and the wrong shape for a bounded model context. Rather than a second result, the layer derives from it — deterministically, in services/agent-observation.ts:
An AgentObservation is never persisted, computes no fact, and duplicates no query logic. Its shared half is identical across all twelve metered warehouse operations:
  • identitystrong / approximate / unresolved / not_applicable, once per observation rather than repeated per row. This is what stops a name resolved by similarity from being reported as a company.
  • filters — the normalized filters that actually ran, in the caller’s own snake_case vocabulary. currency, amount_scope and order_by live here, which is what makes an aggregate total a quantity and a ranking legible as “by count” or “by amount”.
  • facts — what the SERVER computed (an award count, a summed total, a date range). Never recomputed downstream from the rows that happened to fit.
  • visibilitymatched / shown / more_available, clamped so shown can never exceed matched and a trimmed row always sets more_available.
  • coveragestatus / as_of / relations / corpus_rows. complete means complete within the declared corpus AND the filters above; an undeclared or undated corpus yields unknown. corpus_rows sits here, away from the visibility numbers, because beside them it was read as a result count.
  • caveats — the canonical semantic_warnings, as { code, note }. The compact wording lives in services/envelope.ts beside the canonical wording, keyed by the same closed union, so the two cannot drift.
Item projections are per-operation; everything above is shared. The projection is an allowlist, so a field the contract does not name cannot ride through — which is how a search response stays incapable of exhibiting a contact coordinate even if an upstream row carried one. The HTTP surface always applies the abstracted projection. Supplier and opportunity detail are reachable there only through a server-issued record_ref; raw warehouse UUIDs and publisher keys are internal service arguments and are refused over HTTP. Contact coordinates cross the public boundary only through the separately quoted and confirmed Contact Path flow.

Bulk-extraction boundary

Browser endpoints are visible in developer tools by definition, so endpoint obscurity is not a security control. Driftless instead makes the server enforce the boundary: workspace membership is mandatory and removing the creator revokes their API keys; all market-data routes share one 20-request-per-minute actor bucket with a five-minute block; Clerk and OAuth token rotation keeps the same verified human-and-workspace bucket; and the successful-delivery ledger adds a Postgres-backed ceiling of 120 results per rolling hour per actor. Plan allowances still bound calls, credits and page size. A spent idempotency key is refused rather than replaying a newly executed result for free. Cursors and record references are authenticated opaque values; HTTP responses are private, non-cacheable and non-indexable; and the dashboard keeps response data only in memory. Changing IPs, restarting an API instance, replaying raw record ids or asking the browser for a wider projection does not widen access. Which derivation a surface reads is one option on the tool belt: observationProfile: 'full' | 'agent_observation'. Chat SELECTS a derivation; it does not author one. Research stays on full because it cites platform-issued query ids and needs provenance inside the observation itself. CLI and MCP call the services directly and are unaffected. The trace half carries operational metadata only — operation, relation family, rows read vs returned, has_more, identity basis, coverage status, a corpus digest, elapsed ms and warning codes. No name, no amount, no row, no contact. Provenance and evidence stay in the canonical result: a trace explains how the system worked and is never a second place a claim can be sourced from.

Pagination

Keyset, never OFFSET. One uniform PageInfo (limit / returned / hasMore / nextCursor) on every multi-row operation — search, awards/history and awards/aggregate alike, so a caller does not learn a second pagination shape for history or a third for aggregates. Default page 20, maximum 50, limit + 1 fetched so hasMore is a fact rather than a guess. Aggregate groups are a different quantity from search rows — a summarized population, not a raw record — so they default to 50 and cap at 200; one constant (MAX_PAGE_LIMIT / MAX_AGGREGATE_LIMIT in market-data.cursor.ts) is the only place either ceiling is set, so a DTO’s @Max() and the service’s own clamp cannot drift apart. The cursor is opaque and bound to the contract version, the corpus basis, the normalized filters and the query, which produces two refusals with different recoveries: invalid_cursor when the caller changed a filter, cursor_stale when the corpus moved underneath them. A page is bounded in ROWS by limit and, on suppliers/search, optionally in BYTES by max_bytes — a transport bound a caller may set to keep a response inside its own budget, clamped server-side into [1 KiB, 256 KiB] so it can only ever shrink a page. Under it the response carries FEWER rows than limit, with hasMore and the nextCursor that continues exactly where it cut: never a truncated row, never a dropped one. It is not part of what the cursor is bound to, so continuing the same page under a different budget is not a filter change. A limit above the ceiling is still a 400, never a silently smaller page.

Semantic refusals

Every refusal is a DomainException carrying an ErrorCode from ops/error-codes.ts, so it flows through GlobalExceptionFilter with the same code / message / request_id / retryable envelope as every other error. Alongside it travels why, often allowed_values, and usually suggested_correction — because a refusal here is normally a misunderstanding, and the useful answer is the correction. query_too_broad · unknown_state · invalid_observed_kind · invalid_mark_kind · invalid_cursor · cursor_stale · deadline_not_available · amount_requires_currency_and_scope · supplier_name_is_approximate · weak_identity_join · unsupported_aggregation · unknown_filter_value · unknown_facet · amount_scope_not_published · group_by_identity_not_published · text_query_not_searchable · literal_query_too_short · entity_name_not_exact · record_not_available · market_data_timeout · market_data_unavailable · serving_projection_unavailable invalid_mark_kind is search_risks’s equivalent of invalid_observed_kind: mark_kind outside efos | sancion is refused with the allowed pair rather than silently matching nothing. unknown_facet is the same discipline applied to GET capabilities?facets= — a dimension name the capability contract does not compute observed values for is refused with the ones it does, before any query runs. amount_scope_not_published and group_by_identity_not_published are the two ways aggregate_awards can return zero rows for a reason that is NOT “no such market”: rows match every filter but carry none of the requested (currency, amount_scope) pair, or the requested group_by dimension is one no matching row publishes — both are diagnosed by one bounded probe and refused with the pairs or the reason, never served as a silent empty total. unknown_filter_value is the one that earns its keep most often. An exact filter over publisher text is the easiest way in the whole surface to manufacture a lie: contracting_type: "Obra Pública" against a publisher who wrote Obra Publica runs, returns nothing, and reads as a fact about the market. So when a search returns nothing AND an exact categorical filter was applied, the layer checks — in one bounded scan, only on that path — whether the corpus carries those values at all, and refuses with the ones it does carry. A filter that returned rows is self-evidently valid and never pays for the check. No refusal ever contains SQL, a DSN, a token, an internal hostname or a stack.

Matching

Three mechanisms produce a row, ordered by strength, and every result says which one produced it in match.matchMethod: trigram remains in the MatchMethod union for compatibility, but no result carries it: similarity is an INPUT RESOLVER and never a match method. Nothing is retrieved by similarity. The layer chooses, from the query’s shape, and declares the choice. A single token carrying punctuation or digits is a code and takes the literal path; prose takes full text. Entity names never take either: buyer, supplier_name, entity_name and holder_name are RESOLVED to the spellings a publisher used, by normalized-exact match only, and what reaches the database is an equality over those spellings. When no published spelling matches exactly, the layer does NOT fall back to the closest one — it refuses with entity_name_not_exact, hands back the similar names as candidates, and points at the strong identity field where the relation has one. A similar name is frequently a different institution, and its rows are an answer to a question nobody asked. The resolution that did happen is reported in interpretedRequest.normalizations, so a caller can always see which organizations were searched. Risks has no free-text search document at all: a party’s name is either given exactly, resolved to publisher spellings, or not searched — there is no full_text or literal_fallback path into it. Municipality is the one place similarity still resolves silently, and it is narrowly scoped: within an already-exact state, the closest published spelling is chosen, the choice is reported in normalizations, and the filter that reaches the database is an equality against that single value. Municipality follows the same rule and for the same reason: fuzzy geography answers about several towns at once and says so nowhere. Human input resolves to one canonical spelling inside an already-stated state; the filter is exact.

Direct headless clients

The semantic HTTP API is the implementation. CLI and MCP are protocol skins: they preserve the same envelopes and never invoke MarketResearchRunner, a model gateway, SQL, web search or enrichment. ChatGPT, Claude, Codex or a human at the CLI supplies the planning and synthesis. The internal runner remains a separate consumer for Driftless chat and scheduled agent runs. MCP exposes only focused typed market tools. The generic action router and bulk supplier-detail tool are retired; the tools/list budget is an internal 80 KiB latency/context budget, not an external provider limit. Typed operations map to the granular CLI and HTTP surfaces: There are thirteen market tools total. market_capabilities is free discovery; the other twelve are metered. All thirteen carry readOnlyHint: true because they do not modify business data or produce external side effects; the hint does not mean free or exempt from usage accounting. Coverage is not a separate atomic action: every operation envelope already carries the coverage relevant to its answer, while capabilities provides discovery. Adding a duplicate coverage action would consume permanent MCP schema budget without unlocking a new question.

The physical layer

market_data_serving.supplier_search lives in GTM Fábrica (migration 0051). It is derived, disposable and never a source of truth; a build is invisible until complete; publication refuses an incomplete projection, a duplicate observation key, a stale corpus basis or a contact-boundary leak; rollback is one pointer move; and the licence gate is applied on READ. Tenders and awards have no projection, and they no longer need one. The earlier statement here — that the evidence measured them completing and did not justify physical work — was drawn from facet and aggregate queries. It was never drawn from a TEXT SEARCH, because the semantic API did not have one yet. When it did, the shape was pathological: title ILIKE '%q%' OR description ILIKE '%q%' is unindexable, so the planner estimated one row, got ten thousand, and on that estimate re-probed the licence gate once per candidate row. Migration 0054 fixes it in place rather than with a second projection: a stored weighted tsvector on tender_records and award_records, a GIN index, folded trigram and btree indexes on the entity names, and a licence gate evaluated once as an InitPlan instead of once per row. Measured on a synthetic corpus at the cardinalities the frozen profile recorded, for one page of a text search: A projection would have bought a second copy of 112,264 rows, a build pipeline, a publication pointer and a staleness rule, to buy nothing the index does not. Full evidence: gtm-fabrica/docs/market-data/search-plan-evidence.md and scripts/market-data/search-plan-benchmark.js.

Deletion ledger

Legacy audited and PRESERVED, with the reason

What to verify before removing the market-intelligence adapter

One probe settles it. Against the real GTM warehouse, as any role:
If that returns true, the adapter cannot work and the whole MI path (postgres-market-intelligence.adapter.ts, ports/market-intelligence-gateway.port.ts, planning/mastra-market-intelligence-workflow.ts, planning/market-intelligence-tools.ts, planning/market-intelligence-tool-schemas.ts, the MARKET_INTELLIGENCE_GATEWAY binding in radar.module.ts, the three tools in chat-tools.ts, and the two ChatService branches gated on it) is dead and can be removed. Keep planning/capability-bundle.contract.tsexperience-v2/coverage-map.ts type-depends on it.

Manual infrastructure cleanup (not performed here)

This change touches no infrastructure. After merging and deploying:
  1. Delete the Render service GTM Query API Staging.
  2. Delete MARKET_DATA_API_URL from every Driftless environment. The API now refuses to boot if it is still set, so this must happen with the deploy.
  3. Delete MARKET_DATA_API_TOKEN from every environment and from the secret store.
  4. Close the ingress/port opened for the Query API (8931) if one exists.
  5. Keep only the private PostgreSQL connection Driftless uses (GTM_VAULT_DATABASE_URL), its TLS CA and the read-only role.
  6. Confirm market_data_reader can SELECT on market_data_serving.supplier_search and on nothing else in that schema.
  7. Build and publish the first supplier_search projection (node scripts/market-data/build-supplier-search.js --build) — until one is active, supplier search returns serving_projection_unavailable by design.