Skip to content
MAF Learning Hub
Beginner Duration: 20 min

LLM mechanics and prompt fundamentals

Large Language Models (LLMs) are the reasoning engines behind agentic systems. This lesson gives you the mental model you need before wiring agents together in later lessons.

What you will learn

  • How an LLM turns a prompt into a response (tokens, context window, sampling).
  • Why context is finite and how that shapes agent design.
  • The vocabulary used across the rest of Track 1.

How an LLM produces output

An LLM predicts the next token (a chunk of text) given everything currently in its context window. It repeats this, one token at a time, until it produces a stop signal. Two consequences matter for agents:

  • The context window is finite. Everything the model can “see” - system instructions, prior turns, tool definitions, retrieved documents - competes for the same budget.
  • Output is probabilistic. Sampling settings (such as temperature) trade determinism for creativity.

Why this matters for agents

Agents extend an LLM with tools and memory. Because context is finite, agent designers are careful about how many tools and how much data are exposed per turn. This is the “context optimizer” concern: too many tools in one turn degrades reasoning quality.

Key term: context window

The maximum amount of text an LLM can consider at once. Treat it as a scarce budget shared by instructions, history, tool schemas, and retrieved data.

Prompt fundamentals

A well-structured prompt usually separates:

  1. Instructions - the role and task (“You are a support triage agent…”).
  2. Context - retrieved facts or data the model should use.
  3. Input - the specific user request.

Keeping these distinct makes behavior more predictable and is the foundation for the Retrieval-Augmented Generation (RAG) pattern in the next lesson.

Next steps

Continue to RAG and Knowledge Graphs to learn how agents ground their answers in real enterprise data.

References