By late July, two details had to be stated in the present tense: R.A.I.S.E. was actively running on the Z13 reference system, and the retrieval architecture had deepened to eleven evolving stages. Models kept changing. The evidence system became the durable focus.
Why multiple retrieval stages exist
A sensitive evidence question rarely maps cleanly to one fragment. Sources have to be organized, candidate context retrieved, relevance considered, evidence assembled, and the result presented in a form a reviewer can inspect.
The eleven-stage design represented a working multi-stage retrieval and trust pipeline. Evaluation could still combine, split, or replace the internal stages as it exposed failure modes. The number described current architecture, and the company refused to use it as a security rating.
Trust comes from inspection
Extra layers alone do not create trust. The useful property: a reviewer can locate a failure in collection selection, retrieval, context assembly, model behavior, source presentation, or review.
Known-answer tests, visible sources, retrieval-miss analysis, documented exclusions, and human review matter more than a layer count. Retrieval can improve grounding and still return incomplete, irrelevant, or misleading evidence.
Models remain replaceable
The local-model landscape was moving fast, so the product could not depend on one engine. The durable investment was the corpus boundary, retrieval strategy, evidence trail, evaluation process, operating controls, and the Grace orchestration above it.
A clearer model-flexible story started here: use the model the workload and hardware can support, and stop rebuilding institutional memory each time a stronger engine appears.
Research became the public wedge
The July website rebuild shifted the primary story from law firms toward sensitive research. Separate pages explained research, security, air-gapped progress, Grace, the Z13, the pilot, and current build status.
The commercial mechanism became a founder-led 45-day research pilot starting at $7,500, with hardware selected separately. The pilot was framed as a decision instrument: measure retrieval quality, review behavior, failure modes, time saved, and whether the workflow deserved to scale.