Architecture

This section defines the lifecycle and control layers used to build and maintain retrieval-based AI knowledge systems, including Retrieval Augmented Generation (RAG).

Overview

Retrieval-based systems are the pattern where an AI tool answers questions using uploaded documents, policies, contracts, or internal knowledge sources through a retrieval layer.

RAG systems are implemented using different retrieval patterns, including schema-anchored atomic retrieval, segmented document retrieval, and pointer-based document reference. Each pattern introduces distinct control surfaces, failure modes, and governance considerations. This Architecture addresses the structural and lifecycle requirements common to retrieval-based systems, rather than prescribing a single retrieval strategy.

This Architecture focuses on retrieval-based systems because externalizing knowledge into governed sources and retrieval layers introduces structural and lifecycle failure modes that must be managed explicitly. It is not intended as a universal model for all AI use. Related boundary distinctions are documented in Foundations.

It reflects patterns observed while building and maintaining systems over time, where failures often originate from missing structure, unqualified data, weak retrieval boundaries, and unclear ownership. The emphasis is on observed system behavior, including how and why systems drift, degrade, or fail silently after initial deployment.

The architecture is organized as an interdependent lifecycle, not a checklist. Each layer exists to prevent failure conditions that later layers cannot consistently detect or correct. Skipping layers or collapsing controls often produces systems that appear functional early and degrade under sustained use.

Retrieval-based system lifecycle

AI Knowledge System Architecture Lifecycle diagram

High-level lifecycle of a retrieval-based AI knowledge system, showing where structure, boundaries, and governance are enforced.

How this architecture is organized

The architecture is structured as a set of control layers that govern how information is qualified, processed, constrained, and maintained over time.

Preparation layer

Processing layer

Boundary and audit layer

Accountability layer

Synthesis

Scope and constraints

Reading order

The pages are organized in lifecycle order. Each layer defines a control surface that later layers depend on. Skipping layers introduces failure conditions that downstream controls cannot consistently detect or correct.