HMIS-NL, verified AI analytics
I built HMIS-NL so homeless-services staff can ask outcome questions in plain English and put the answer in front of a funder. The model reads the question. Two engines compute the number, and scheduled agents prepare the reports that come due.
Hosted web app and Textual terminal client. Not publicly linked yet.
- Hosted app
- Public sample data
- Private source
- Problem
- Funders and boards expect outcome numbers an agency can defend. A language model will produce a confident number that nobody can trace.
- Engineering decision
- The model only chooses a typed request through tool calls. Two engines compute the number and a validator blocks any answer that fails. Scheduled agents handle the recurring reporting work.
- Result
- Outcome numbers checked by two engines before anyone sees them
- My role
- Sole engineer, independent analytics application
- Source access
- Private repository. Live demo uses public HUD sample data.
In this case study Features and engineering details
What it does
- Two engines per answer
- A second engine recomputes each number from a declarative spec. If the two disagree, the answer is withheld.
- Deadline agent
- Every morning it prepares a packet for each funder report that is overdue or due within 45 days.
- Router that improves overnight
- Questions that failed to route are retried with rephrasings. Versions that pass wait for a person to approve them.
- Signed proof receipts
- Each answer's receipt is hashed and signed, so a figure in a board packet can be checked later.
From question to checked answer
Typed request
The model calls a tool to name the metric and period, or to ask a clarifying question.
Two engines
Python computes the answer. A second engine recomputes it from the metric's declarative spec.
Validator gate
Ordered checks re-resolve the rule from the registry. A failure replans or refuses.
Answer from evidence
The sentence staff read is built from the evidence panel. The model never writes it.
Walkthrough
Recorded on September 24 and 25, 2026 against the hosted demo, which runs on public HUD test data with no real client records yet. The walkthrough moves from counting people served to outcomes, then shows clarification, spoken questions, board reports, the By-Name List, the deadline agent, proof receipts, HUD citations and the terminal client. Narrated with a synthetic voice. Captions are available in the player.
Keep the model away from the numbers
A reporting question leaves room for interpretation, and a language model will fill that room with a confident number. I limited the model to one job. It reads the question and calls a tool with a typed request for one metric over one period.
That request runs through a bounded state machine. If the validator rejects it, the system applies the validator's fix hints and replans up to twice with no model call. It can also give the model one repair turn that includes the engine's exact rejection. After that the run answers or refuses.
The sentence staff read is built from the evidence panel. The model never writes answer text, so it has no way to slip a number into it.
Check every number twice
Each metric has a declarative spec of who counts and how. A second engine interprets that spec in plain Python and recomputes every answer the primary engine returns. If the two disagree, the answer is withheld and the refusal names the metric and period that failed.
The labels stay honest. A metric with no spec is marked single engine. If the second engine reuses the first engine's population code, the answer is not marked verified, because that check would not be independent.
A validator also runs ordered checks on the request itself. It re-resolves the rule from the registry instead of trusting the model's draft. No path shows a number after a failed check.
Exits to permanent housing in 2022, checked by both engines. The panel shows the counting rule and the project types included. Public HUD test data.
View full imageAsk before guessing
Some questions cannot be answered until someone picks a meaning. "How many people did we serve?" needs a period, and the choice changes who is counted. The model can call a clarify tool, and fixed rules catch cases like a missing period or a follow-up with nothing before it.
The server holds the original question while staff pick an option. Each run can clarify once. After that it has to answer or refuse, so it never loops on questions.
The hosted demo asks for a period before counting. The current reporting quarter is suggested. Recorded September 25, 2026.
View full imageAgents that run before staff arrive
Two scheduled agents run every morning on the demo server. The deadline agent reads the contracts calendar and finds each funder report that is overdue or due within 45 days. For each one it runs the engine over that grant's own window and writes a packet. The packet holds the numbers, how they were checked, what could not be produced and what a person still has to do. It replaces the last packet for that report and files it with the agency's documents, where the policy assistant can cite it.
The router agent works on questions the system could not answer. It retries each one with deterministic rephrasings. When a rephrasing passes the validator, it becomes a proposed example for the router, and a person decides whether to promote it. A daily sweep asks paraphrases of approved questions to catch answers that drift. A weekly job checks HUD's published standards for changes and re-ingests any that moved.
No agent writes a number. Every figure in a packet comes from the deterministic engine.
The contracts calendar the deadline agent reads each morning. Award and drawdown figures are sample finance data. Recorded September 24, 2026.
View full imageExplain why a number moved
HMIS data keeps changing after a quarter closes. Late entries and corrections can move a number that already went to a funder. The restatement report reruns the headline metrics against the previous export and the current one.
A third run uses the current export limited to enrollments that existed at the previous export date. That splits each change in two. One part comes from enrollments created since the last export. The other comes from edits to enrollments that were already there. The report describes the change and leaves the cause to staff.
October to December 2021 against the export of January 14, 2022, on public HUD test data. Recorded September 25, 2026.
View full imageProof receipts and small-count privacy
Every answer carries a proof receipt. It records the question, the rule and its version, the period, the export fingerprint and the value. A SHA-256 hash covers the record and an HMAC with the deployment key signs it. Anyone holding a number from a board packet can check it later against the same export and rule set.
Small counts can identify people. Counts below the threshold are masked at the single point where results are computed, before any hash is taken, so a receipt cannot be used to recover them. Tables and CSV exports keep them labeled as hidden, and charts leave them out. An MCP server exposes the same metrics to other tools with a disclosure ledger, so a second call cannot subtract out a hidden count.
Two groups stay hidden under the threshold of 11. Zero and hidden counts are displayed differently. Public HUD test data.
View full imageCite the rule in HUD's own words
Staff also need the rules behind a number. HUD standards and an agency's own policies live in separate assistants, so the sources never mix. Codes and field names go to an exact lookup. Definitions go to hybrid retrieval, which fuses keyword and embedding search.
Answers must cite their sources or say the documents do not cover the question. Any number in a policy answer has to appear word for word in a cited passage. A citation opens the original PDF at the right page with the passage highlighted, using word positions extracted from the PDF.
A HUD Standards answer opens the HMIS Data Dictionary at the cited passage on page 11.
View full image