Applied AI / Built Environment / 2026

Josiah Lau

Applied AI for the decisions that shape buildings and construction.

I connect source-grounded knowledge, approved requirements, constrained geometry, traceable quantities, and simulated action in systems that remain inspectable from input to evidence.

DomainDesign + construction
Primary workRAG + embodied AI + computational design
EvidenceRunnable code + evals + visible limits

Generated portfolio concept image / not hardware evidence

Selected Work

Three projects, three kinds of engineering evidence

The strongest work combines implementation depth, measured behavior, visual explanation, and explicit failure boundaries. Metrics come from fixed public-source snapshots, bundled synthetic fixtures, or simulator holdouts as labeled beside each result.

Local AEC retrieval interface with a grounded answer and source citations 01 / Flagship

AEC Code Compliance RAG

Public-source ingestion, four retrieval modes, citations, abstention, evals, and focused tests.

Document retrieval is stronger than exact evidence retrieval Document Hit@1 0.952 / MRR 0.976; exact-target Hit@1 0.810 / MRR 0.881; page Hit@1 0.778 on manually labeled fixed-snapshot targets. Public Hit@1 95% interval [0.773, 0.992], n=21. No-answer is 2/2; this is not compliance certification or broad accuracy evidence. Read visual case study
Construction-grid observations and measured holdout behavior across state and rendered-pixel policies 02 / Embodied AI

Construction Embodied Agent Simulator

Shared expert demonstrations, rendered-pixel shift tests, raw and filtered rollouts, and planar MuJoCo command replay.

0.760 filtered success / 51 raw contacts / 0 filtered contacts 96 unseen grids for policy evaluation; 12-scenario physics subset with named rigid-contact telemetry. Rendered-RGB success falls from 0.719 to 0.427 under an unseen palette. Pixels come from state; no camera or hardware evidence. Read visual case study
Constraint-aware massing plan, isometric view, and transparent proxy measurements 03 / Computational Design

Constraint-Aware Massing Explorer

Seeded geometric options, hard constraints, Pareto ranking, editable objectives, and transparent environmental and access proxies.

0.977 feasible rate / 0.19% mean best GFA error 864 candidates per method across synthetic sites; unconstrained baseline feasible rate 0.052 and mean best GFA error 34.47%. Rectangular envelopes and transparent proxies only; no code inference, calibrated environmental simulation, or approvable design. Read visual case study
Measured target-reach and rigid-contact rates for raw, filtered, and reference embodied-agent command traces
Headless planar MuJoCo replay: raw egocentric commands contact named rigid geometry on 51/150 movements; filtered egocentric and A* traces record 0 contacts across 148 and 98 movements. Not a mobile-robot or safety validation.

One-Command Technical Check

python scripts/reviewer_check.py

Checks public claims, documentation commands, links, app titles, and focused tests for the three primary projects without rewriting tracked evidence.

View the artifact ledger

Executed AEC Integration

Approved requirements to a bounded schematic cost input

A deterministic local runner connects the specification, massing, and QS interfaces. Every conversation, site value, generated option, and rate in this trace is synthetic.

Executed AEC system journey with two input streams, source and approval gates, typed project handoffs, traceability, and professional review
5 approved / 3 mapped / 2 retained; 16/16 sourced fields; 96 candidates / 92 feasible / 42 storey-matched; 7/7 priced takeoff lines / tender not_run.

Budget and accessibility remain outside automation. The QS handoff covers one schematic ground-floor envelope, not a developed plan or whole-building estimate.

python integrations/aec-design-to-cost/run_workflow.py Inspect contract, tests, and trace
Role-tagged project messages, requirement ledger, and source-linked draft clauses Project Communication

Communication and Specification Assistant

Shared role-tagged chat, versioned requirements, conflict detection, scoped approvals, audit events, and source-linked draft clauses.

33-case language stress: 0.958 F1 / 0.939 exact 5 conversations / 35 messages / 1.000 fixture F1; 2 number-word misses are retained. Manually labeled synthetic cases, not open-domain chat or a professional specification. Read visual case study
Constraint-aware massing plan, isometric option, and proxy metrics Computational Design

Constraint-Aware Massing Explorer

Hard geometry checks, baseline comparison, Pareto ranking, and inspectable proxy scores.

No code inference, internal egress, calibrated simulation, structure, or professional design output. Read visual case study
Vector quantity takeoff, cost build-up, and synthetic tender comparison QS Automation

Takeoff and Tender Analysis Workbench

Shared-wall measurement, opening deductions, rate provenance, uncertainty bands, and line-level tender exceptions.

21/21 quantities / 5/5 tender exceptions No PDF/CAD/BIM parsing, live pricing, professional QS output, or award recommendation. Read visual case study

Capabilities

Domain reasoning backed by inspectable engineering

Architecture contributes the problem framing: constraints, alternatives, coordination, review gates, and accountable handoff. Software and ML evidence still come from code, fixed datasets, tests, and reproducible artifacts.

01

Source-grounded AI

Page-aware ingestion, metadata-rich chunks, retrieval baselines, citations, abstention, and evaluation.

02

Embodied AI

Observation encodings, behavior cloning, appearance shift, action filtering, and planar physics replay.

03

Computational design

Parametric geometry, hard-constraint validation, Pareto ranking, and explicit proxy objectives.

04

AEC workflow systems

Requirements, approvals, source lineage, quantity provenance, cost build-up, and human review boundaries.

05

Engineering practice

Typed APIs, SQLite state, automated tests, generated evidence, idempotence checks, and local-first delivery.

Open the evidence-linked skills matrix

Primary Project

AEC Code Compliance RAG

This is the most developed project in the portfolio. It applies built-environment domain knowledge to Markdown and page-aware PDF ingestion, Singapore public-source manifests, source filters, TF-IDF, BM25, dense LSA, hybrid retrieval, citation checks, unsupported-scope handling, and a tested local service with bounded durable telemetry.

Evaluation
51 synthetic cases plus 24 authored public-source cases over a fingerprinted 15-document snapshot
Synthetic check
51-case synthetic set: Recall@4 1.000 / MRR 0.906 / Hit@3 1.000
Evidence location
Document Hit@1 0.952 versus exact manually labeled target Hit@1 0.810; source-page Hit@1 0.778
Uncertainty
Public Hit@1 0.952 [0.773, 0.992], n=21; MRR 0.976 [0.929, 1.000]. Hybrid-vs-BM25 MRR delta 0.012 [0.000, 0.036], inconclusive
Service
12/12 local contract checks passed; 9 requests observed before metrics response. The fixed workload returned 48/48 HTTP 200 responses, 0 server errors, and P95 at or below 500 ms
Persistence
48 payload-free query telemetry rows remained after app reconstruction
Proof
Architecture, eval design, paired uncertainty, ablation, failure analysis, demo answers, service reports, and tests
Limit
No authority validation, professional sign-off, OCR, network load, uptime, external deployment, traffic, or live amendment guarantee
Local AEC RAG Streamlit interface showing an answer and source citations
Actual local Streamlit run using labeled demo data.
Document-level, exact-chunk, and source-page retrieval metrics for the fixed AEC public-source snapshot
Document Hit@1 0.952 / MRR 0.976; exact-target Hit@1 0.810 / MRR 0.881; source-page Hit@1 0.778 / MRR 0.861. Finding the expected publication is easier than retrieving the labeled supporting passage. Exact targets were authored from the fixed local snapshot and have not been independently reviewed.
Point estimates and 95 percent intervals for the authored AEC public-source retrieval snapshot
Public Hit@1 0.952 has a 95% Wilson interval of [0.773, 0.992] over 21 answerable cases; MRR 0.976 has a 95% bootstrap interval of [0.929, 1.000]. No-answer 1.000 is 2/2 with a Wilson interval of [0.342, 1.000]. The intervals resample this fixed authored set; they are not external-validity claims. Exact-target uncertainty: Hit@1 0.810 [0.600, 0.923] and MRR 0.881 [0.762, 0.976] across 21 manually labeled targets.
Synthetic construction grid, rendered RGB crops, and measured holdout results across four observation families
Generated from evaluator output. RGB pixels are rendered from privileged simulator state, not a physical camera.
Concept illustration of a construction robot following a planned safe route
Generated concept image. It is not a simulator screenshot or hardware evidence.

Embodied AI

Language, state, action, and visible failure

The simulator maps task instructions and structured construction grids to closed-loop actions. It compares engineered-state, semantic-raster, agent-centered local-state, and rendered-RGB imitation on the same demonstrations and holdout layouts. An unseen palette exposes visual brittleness.

Policies
Random forest plus world-raster, egocentric-state, and rendered-RGB MLPs fitted from shared A* demonstrations
Holdout
96 disjoint procedural scenarios with raw and filtered closed-loop rollouts
Finding
Egocentric action accuracy 0.834 versus world-raster 0.478; filtered success 0.760 with 943 interventions
Pixel ablation
Rendered-RGB action accuracy falls from 0.813 to 0.474 when pixels are replaced by training means
RGB shift
Action accuracy 0.813 to 0.417; raw success 0.635 to 0.000; shifted filter interventions 3315
Physics replay
12 unseen scenarios; raw contact rate 0.340 versus filtered 0.000 in a planar MuJoCo command adapter
Limit
Pixels are rendered from privileged state and the filter has full rules; physics uses a planar body and static proxies, not a physical camera, mobile robot, ROS, hardware, or safety validation

Architecture Background

Design discipline carried into AI engineering

Architecture provides the domain context for the AEC and embodied-AI work: requirements, constrained alternatives, versioned information, human review, and accountable handoff. It does not replace software evaluation or robotics validation.

Read the project statuses and image provenance
Architecture-to-AI transfer map

Professional Delivery

Experience across design, coordination, tender, and handover

The AI work is informed by direct responsibility for drawings, project information, stakeholder coordination, procurement, progress, defects, and delivery. These roles provide domain context; the repository remains the evidence for software and ML claims.

  1. 2024 - presentDirector / 3JAI StudioAI workflow systems, review criteria, and client delivery.
  2. 2022 - 2024Project Manager / People's AssociationTender, cost, progress, defects, audit, and handover.
  3. 2020 - 2021Architect / HDBGIS, estate planning, environmental studies, and agency coordination.
  4. 2020Graduate Architect / aKTa-RCHITECTSAuthority submissions and construction packages.
  5. 2018 - 2019M.Arch Intern / Tierra DesignGrasshopper CAD automation and drawing QA.
Singapore M.Arch MSc Design + AI Revit + Rhino/Grasshopper + Python

Project Scope

What this portfolio does and does not show

Implementation: runnable code, tests, fitted classical models, persistence, metrics, and simulation.

Data: project pages identify synthetic conversations, sites, rates, tenders, public subsets, and locally downloaded sources.

Mock providers: deterministic LLM and VLM substitutes test interface contracts, not model quality.

Simulation: robotics results do not imply hardware, perception, or physical safety validation.

Professional work: no customer adoption, deployed reliability, compliance outcome, QS sign-off, tender advice, or approvable design is claimed.

Contact + Evidence

Continue into code, evaluations, and project context