Research at Spatial Web is applied: it exists because the corpus exists. Five operating companies produce photoreal, measured, labeled interiors every day — the program studies what machines can learn from them. Papers publish here first; drafts circulate earlier by request.
What listing capture actually contains: photoreal imagery, lidar-derived geometry, listing-derived semantics, and verified market records — assembled as a by-product of operating real estate companies. Covers scale, formats, provenance, and rights, and argues that lived-in homes are the missing distribution in physical-AI training data.
What changes when language models are grounded in the built environment: tokenizing space, aligning geometry with description, and evaluation beyond captioning. Sets out the architecture and data requirements — and the case that market-grade metadata outperforms crowd labels.
Market records as free ground truth: using ownership, valuation, and listing metadata to supervise indoor scene understanding without an annotation pass. Every home arrives pre-labeled by the transaction that sold it.
One capture, four layers. Each environment in the corpus carries all four — produced by commerce, joined at the address.
Fig. 1 — spec sheet, abridged. Full specification ships with SW-26.01.