The Three AI System Approaches
Definitional taxonomy of the three system level approaches used to control AI behavior: prompt engineering, fine tuning, and Retrieval Augmented Generation (RAG).
Purpose
AI systems that appear similar at the interface level often behave very differently in practice. Those differences are not primarily caused by the model itself, but by how the system is designed to influence model behavior. This document defines the three system level approaches used to control AI behavior and clarifies how they differ in structure, scope, and consequence.
This page is intentionally definitional. It does not recommend tools, vendors, or architectures. It exists to establish a stable taxonomy that downstream Architecture, Foundations, and Analysis content can rely on without redefinition.
Introduction
Modern AI systems use three distinct approaches to control model behavior: prompt engineering, fine tuning, and Retrieval Augmented Generation. Each approach is defined by what is modified in order to influence outputs: the input supplied to the model, the model itself, or the context provided at inference time.
These approaches are peers. They are not technologies, frameworks, or architectures. They describe fundamental strategies for behavior control. Practical AI systems may combine them, but the distinctions remain necessary to reason accurately about behavior, failure modes, and governance.
Prompt engineering
Prompt engineering is an AI system approach in which model behavior is shaped exclusively through instructions, constraints, examples, or contextual clarification provided at runtime. The underlying model is not altered, and no external knowledge is selectively retrieved. All influence occurs through the prompt itself.
In a prompt engineered system, the model relies entirely on its existing training and the phrasing of the request. Behavior changes are temporary and session bound. If the prompt is removed or rewritten, the behavior disappears. This makes prompt engineering flexible and immediate, but structurally fragile under complexity or ambiguity.
Prompt engineered systems may include uploaded documents, project level instructions, or rule files. In these systems, documents influence output by being interpreted directly by the model rather than by selectively constraining what the model is allowed to see. When documents or rules are supplied wholesale and interpreted at generation time, the system remains prompt engineered even if files are indexed or summarized internally. Selective retrieval that controls model access at inference time is what distinguishes Retrieval Augmented Generation from prompt engineering.
Prompt engineering is widely used by individuals, analysts, developers, and small teams because it requires no infrastructure and produces immediate results. It is common in general purpose conversational use, exploratory analysis, drafting, and early experimentation.
The structural limitation of prompt engineering is that it cannot add new knowledge, enforce correctness, or guarantee consistency. The model may comply with instructions, reinterpret them, or ignore them while sounding confident. Improved wording is often mistaken for reliability.
Fine tuning
Fine tuning is an AI system approach in which a pre trained model’s internal weights are modified through additional training on a specialized dataset. Instead of relying on runtime instructions, fine tuning permanently biases how the model behaves.
In a fine tuned system, behavior changes persist across prompts and sessions. The model begins to exhibit domain specific patterns even when given minimal instructions. This makes fine tuning effective for narrow, repetitive tasks where output consistency matters more than adaptability.
Fine tuning is commonly used in enterprise and platform settings where tasks are stable and high volume, such as classification, labeling, templated generation, or constrained customer interaction. Its defining tradeoff is that assumptions and errors become embedded directly into the model.
The primary cost of fine tuning is operational rather than financial. Training data must be representative, accurate, and continuously maintained. Small or poorly curated datasets frequently degrade output quality compared to well trained base models accessed through APIs. Mistakes persist until retraining occurs, making fine tuning poorly suited to rapidly changing or contested knowledge.
Retrieval Augmented Generation
Retrieval Augmented Generation, abbreviated as RAG, is an AI system approach in which model output is constrained by external source material retrieved at inference time. When a query is made, relevant content is selectively retrieved from an external corpus and supplied to the model as context for generation. The base model’s weights remain unchanged. What the model is allowed to see is determined by retrieval.
RAG is a broad approach. Systems described as RAG vary widely in behavior, reliability, and governance depending on where interpretation is allowed to occur, including segmented document retrieval and pointer based document reference patterns.
Valid but distinct RAG operating modes
Both modes meet the definition of RAG. They differ in where interpretation occurs and how much failure the system is designed to tolerate.
Chunked and vectorized documents retrieved at query time are a valid form of RAG. This pattern is widely used by AI consulting teams, internal IT or data teams, municipalities, and small and mid sized organizations where AI is expected to assist with locating, summarizing, or navigating source material rather than acting as an authority.
In these generation first RAG systems:
- Documents are split into chunks and indexed, often using vector similarity.
- Retrieval supplies relevant text fragments.
- Interpretation occurs at query time inside the language model.
- Humans remain the authority and are expected to read or verify retrieved material.
- Silent failure and ambiguity are tolerated because outputs are advisory.
This operating mode is appropriate in environments where convenience and speed outweigh enforceable correctness, and where AI outputs are not reused as durable knowledge.
A second RAG operating mode exists for different requirements.
In knowledge constructed RAG systems:
- Source material is interpreted upstream.
- Atomic facts or claims are extracted and qualified before retrieval.
- Knowledge units are bound to explicit schema, scope, and authority rules.
- Retrieval boundaries are enforced before generation.
- Authoritative interpretation does not occur at query time.
This mode appears in systems operated by internal AI teams or consulting teams where outputs are reused, relied upon over time, or expected to behave consistently. It exists specifically to support traceability, verification enforcement, and visible failure when required knowledge is missing.
Both operating modes are RAG. They differ not by maturity or tooling, but by interpretation ordering and tolerance for failure.
Where complexity lives
In all RAG systems, primary complexity lives upstream. Data qualification, extraction discipline, knowledge modeling, and boundary definition determine system behavior. Retrieval algorithms and prompt design cannot compensate for weakly defined source material.
RAG exists to allow knowledge to live outside the model and to change without retraining. Whether that knowledge can be governed, verified, or enforced depends entirely on whether interpretation is allowed to occur during generation or is constrained before retrieval.