bimq
Read-only BIM query server for agents: structured, policy-bounded queries on IFC/gbXML models with deterministic cited results.
README
bimq
A read-only BIM query server for agents. Point it at an IFC or gbXML model and it answers structured questions — fire-rated doors on level 3, elements with no material assigned, spaces below the minimum daylight area — bounded by a policy file, with deterministic results and a citation back to the GlobalId and source line behind every row.
Zero dependencies. Python 3.11+. MCP server over stdio, plus a CLI that answers the same questions so you can check a policy before you trust an agent to it.
bimq query elements_by_property model.ifc \
type=IfcDoor storey="Level 3" property=FireRating op=exists
id type name tag storey_name source match
---------------------- ------- -------- ---- ----------- -------------- -------------------------------
0XBbD$nZDLuRru91_CQ_xe IfcDoor Door-302 D302 Level 3 office.ifc:314 Pset_DoorCommon.FireRating=EI60
31kamnSrrNbf3eF0_vhXrJ IfcDoor Door-301 D301 Level 3 office.ifc:304 Pset_DoorCommon.FireRating=EI60
2 row(s)
digest: sha256:06321f331735417dd149d649b8e26de71a63ff2bb26a31c38cbf66a4f0314b77
Then check it, because a citation you cannot follow is just a confident-looking string:
bimq cite model.ifc 31kamnSrrNbf3eF0_vhXrJ
31kamnSrrNbf3eF0_vhXrJ (IfcGloballyUniqueId)
office.ifc:304 #297
#297= IFCDOOR('31kamnSrrNbf3eF0_vhXrJ',#5,'Door-301',$,$,$,$,'D301',2100.0,900.0,.DOOR.,.SINGLE_SWING_LEFT.,$);
Why
The current instinct is to dump IFC text into a context window. That fails immediately at real model sizes, and it fails quietly: a 300 MB model is roughly 95% geometry, so what fits in the window is a truncated arbitrary slice, and the model answers from it anyway. The failure looks like a fluent paragraph about a door that does not exist.
bimq inverts it. The model stays on disk. Queries are structured, the answers are small, and every row carries the id and line it came from — so a claim can be checked against the file instead of trusted.
Three properties hold for every answer:
Bounded. A TOML policy file says what is readable — which files, which queries, which types, which storeys, which properties. The engine reduces the model to the visible set before the query runs, so a query primitive cannot reach what the policy hides even by accident.
Deterministic. Same model, same query, same bytes. Every answer carries a
digest you can pin in a test. Element order, group order and float rounding are
all fixed; the read block size and the file's name do not change a finding.
Cited. Every row carries {id, id_kind, source, ref, line}. bimq cite
resolves it back to the original statement. The test-suite re-reads the recorded
line for every element of every fixture and fails if the id is not there.
Install
pip install bimq
Or run it from a clone with no install at all — there is nothing to build:
python -m bimq describe tests/fixtures/office.ifc
Use it as an MCP server
{
"mcpServers": {
"bimq": {
"command": "bimq",
"args": ["serve", "/srv/bim/tower.ifc", "-p", "/srv/bim/policy.toml"]
}
}
}
Every query primitive becomes a bim_* tool, all annotated readOnlyHint, plus
bim_cite. Omit the model path to let each call name its own file — then
allow_sources is what stands between a path argument and your filesystem.
The server tells the agent how to behave on initialize: call bim_model_summary
first, quote a GlobalId for anything you assert, treat truncated as "there are
more", and read notes — because no results and no data recorded are
different findings and only the notes distinguish them.
A policy refusal comes back as a successful tool result carrying
policy_denied and the rule that fired, not as a protocol error. An agent that
receives a protocol error retries; an agent told "this policy does not expose
costs" reports the limit and moves on.
Query primitives
| Primitive | Answers |
|---|---|
model_summary |
schema, units, storeys, entity types, property-set names |
spatial_tree |
project → site → building → storey → space |
elements_by_type |
elements of a type, subtypes included |
elements_by_property |
property comparison; the fire-door workhorse |
elements_missing_property |
data completeness: who has no value for this field |
elements_missing_material |
no material through any of IFC's five ways of saying so |
spaces_by_area |
rooms inside an area range, always in m² |
property_values |
distinct values with counts — run this before guessing names |
quantity_rollup |
totals grouped by type, storey or PredefinedType |
element_detail |
expand specific GlobalIds to every pset and quantity |
bimq queries prints their parameters. List queries return compact rows on
purpose; element_detail is the drill-down, and keeping those separate is what
stops a query from becoming the context dump it replaced.
Aggregates report their own coverage. A roll-up over 200 walls where 160 carry no
quantity says so in summary and notes, because a total over 40 of 200 is not
a total.
Policy
name = "consultant-readonly"
[allow_sources]
roots = ["/srv/bim"]
max_bytes = 536870912
[allow_queries]
queries = ["model_summary", "spatial_tree", "elements_by_type", "elements_by_property"]
[scope_storeys]
names = ["Level 2", "Level 3"]
include_unplaced = false
[allow_types]
types = ["IfcBuiltElement", "IfcSpace", "IfcBuildingStorey"]
[deny_properties]
properties = ["*Cost*", "Pset_Tender.*"]
[redact_properties]
properties = ["*.Owner*", "*SerialNumber*"]
placeholder = "[redacted]"
[max_results]
limit = 200
bimq policy check policy.toml # validate before shipping
bimq rules # every rule, with an example
Notes on the design:
- An unknown table is a hard error, not a warning. A file whose job is to withhold data must not fail open because of a typo.
denyandredactare different tools. A denied property is gone; a redacted one is present with a placeholder. The distinction matters to an agent: redaction says this exists and you are not being shown it, so the agent reports a gap instead of concluding nobody entered the data.- Denial covers the query side too. You cannot filter on a denied property,
because
op=gt value=1000repeated a few times reconstructs it. - Withholding is reported, never silent. Answers carry
policy.elements_withheldand a note. Truncation setstruncated: true. - Every answer is capped even with no policy at all. "Unlimited" is not a sane default for something feeding a context window.
Source formats
| Format | Notes |
|---|---|
IFC-SPF (.ifc, .ifczip) |
IFC2X3 / IFC4 / IFC4X3, streaming reader, no dependencies |
gbXML (.gbxml) |
energy models; ids are stamped gbXMLId, never confused with GlobalIds |
Wanted, one per PR: Revit export (pyRevit/Dynamo JSON), Speckle stream, IFC-JSON, COBie. See CONTRIBUTING.md.
How the IFC reader stays small
bimq/sources/spf.py is a complete ISO 10303-21 reader in under 400 lines. The
parts that matter:
- The file is scanned in 4 MB blocks, so a 300 MB model is never one string. A block boundary can land inside a string literal, so the scanner explicitly matches unterminated literals and carries them forward. Tested at block sizes down to one byte, where the result must still be byte-identical.
- A
;inside'a;b'does not end a statement,''is an escaped quote, and\X2\...\X0\decodes to UTF-16 — soPhòng họpsurvives the round trip. - Comments appear between statements, inside parameter lists, and around section markers. All three are handled; the reported line still points at the entity.
- Geometry is never loaded. An instance is kept only if its first attribute
is a syntactically valid GlobalId — making it an
IfcRootsubtype — or if it is one of ~30 unrooted carriers of property, quantity, material or unit data. The test is applied to the raw text before tokenising, which is where the parse time on a real file actually goes.
Check the throughput claim yourself without needing a model of your own — this writes a file shaped like a real export (a modest element count buried in geometry), parses it, and reports:
$ bimq bench --synthetic 20000
synthetic model: 20000 elements among 820006 instances
tmp6l05ix9x.ifc: 32.2 MiB, 20001 elements
parse: 1.71 s · 18.8 MiB/s · 11,680 elements/s
peak rss: 107 MiB (3.3x file size)
820,006 instances go in; 20,001 elements stay resident. That ratio is the whole argument — resident size tracks how many things the building has, not how many points were needed to draw them.
Model files are treated as untrusted input. The gbXML reader refuses entity
declarations outright, so a file cannot carry a billion-laughs expansion or an
external entity pointing at /etc/passwd.
Units
Every number bimq returns is SI: metres, m², m³. A model authored in millimetres
with areas in square metres (what Revit exports) and one authored in feet with
areas in square feet both answer spaces_by_area max_m2=8 correctly.
IfcConversionBasedUnit chains are resolved, not guessed.
Try it
The fixtures are synthetic — stated plainly, because a fixture pretending to be a
real project is one nobody can check. What makes them useful is that the defects
are deliberate and enumerated: a wall with no material, a fire door with no
rating, a room below 8 m², a room with no area quantity at all, a door whose
rating is inherited from its type, and a Vietnamese room name written with \X2\
escapes.
python scripts/make_fixture.py tests/fixtures
bimq describe tests/fixtures/office.ifc
bimq query elements_missing_material tests/fixtures/office.ifc type=IfcWall
bimq query spaces_by_area tests/fixtures/office.ifc max_m2=8
bimq query property_values tests/fixtures/office.ifc type=IfcDoor property=FireRating
bimq query spaces_by_area tests/fixtures/legacy-imperial.ifc max_m2=8 # authored in feet
bimq query spaces_by_area tests/fixtures/clinic.gbxml max_m2=8 # gbXML, same primitive
Development
make test # the suite
make fixtures # regenerate fixtures (byte-identical; CI checks this)
make bench # parse throughput on a fixture
License
MIT
Recommended Servers
playwright-mcp
A Model Context Protocol server that enables LLMs to interact with web pages through structured accessibility snapshots without requiring vision models or screenshots.
Magic Component Platform (MCP)
An AI-powered tool that generates modern UI components from natural language descriptions, integrating with popular IDEs to streamline UI development workflow.
Audiense Insights MCP Server
Enables interaction with Audiense Insights accounts via the Model Context Protocol, facilitating the extraction and analysis of marketing insights and audience data including demographics, behavior, and influencer engagement.
VeyraX MCP
Single MCP tool to connect all your favorite tools: Gmail, Calendar and 40 more.
graphlit-mcp-server
The Model Context Protocol (MCP) Server enables integration between MCP clients and the Graphlit service. Ingest anything from Slack to Gmail to podcast feeds, in addition to web crawling, into a Graphlit project - and then retrieve relevant contents from the MCP client.
Kagi MCP Server
An MCP server that integrates Kagi search capabilities with Claude AI, enabling Claude to perform real-time web searches when answering questions that require up-to-date information.
E2B
Using MCP to run code via e2b.
Neon Database
MCP server for interacting with Neon Management API and databases
Exa Search
A Model Context Protocol (MCP) server lets AI assistants like Claude use the Exa AI Search API for web searches. This setup allows AI models to get real-time web information in a safe and controlled way.
Qdrant Server
This repository is an example of how to create a MCP server for Qdrant, a vector search engine.