Part 2 in the Demystifying Recent AI Advances Series

In Part 1 of this series, we toured the model zoo. We covered what separates LLMs, SLMs, VLMs, reasoning models, and diffusion models, and the deeper split between discriminative and generative approaches. Knowing which model to reach for is a good start. But a model on its own is inert. It answers when spoken to, and that is the whole of the interaction.

This post covers the two layers that sit on top of the model and largely determine what an AI system can actually accomplish. The first is autonomy: how much initiative the system takes, what it is permitted to touch, and how much human oversight it requires. The second is the plumbing: the standards and structures, namely Skills, the Model Context Protocol, and the agent harness, that let a model reach beyond its own context window and take action in the world.

Both layers are where much of the recent terminology explosion has landed, and both are where the distinctions matter most in practice.

The Autonomy Spectrum: Chatbots, Copilots, and Agents

Another area ripe for confusion is the spectrum of AI autonomy. The terms chatbot, copilot, and agent get used loosely, but they describe meaningfully different levels of capability and independence.

Chatbots are the most familiar. They are reactive, conversational interfaces that wait for a prompt and generate a text response. No tool access, no persistent state beyond the current conversation window, no ability to take action in the world. Ask a chatbot a question, and it answers. That is the extent of the interaction. Most consumer-facing AI assistants still operate at this level.

Copilots represent a step up. A copilot augments human work by suggesting actions: code completions, document edits, analytical next steps, dashboard configurations. The critical distinction is that the human remains in control. The model proposes; the human takes action. GitHub Copilot suggests code completions, SAS Viya Copilot recommends model pipeline configurations, and an AI assistant drafting an email for your review are all copilot-level interactions. Copilots accelerate human productivity without removing human judgment or approval from the process.

Agents are where the real paradigm shift occurs. An agent operates autonomously in a loop: it plans, takes action by invoking tools, observes the results, and decides what to do next. The critical differentiator between a copilot and an agent is autonomy plus tools. Instead of proposing the next step for a human to accept, an agent executes its own plan against external systems. That said, autonomy does not mean unchecked capabilities. It is often wise and prudent to design agents so that any potentially destructive or irreversible actions, such as deleting data, sending communications, modifying production systems, or committing financial transactions, are gated behind explicit human review and approval. Only after an agent has proven itself reliable should these gates be relaxed, and even then, only in a very controlled and observable way.

The best agent architectures make this easy by design, enabling routine read-only and low-risk operations to flow freely while surfacing high-stakes decisions to a human for confirmation. The goal is not to hobble the agent but to ensure that the speed and scale of automation never outrun human judgment on the actions that matter most.

To paraphrase Andrej Karpathy, from his YouTube clip, “[With AI] we can outsource intelligence, but we cannot outsource our understanding or responsibility.” Table 1 compares and contrasts these tools.

Dimension Chatbot Copilot Agent
Interaction Reactive: waits for a prompt Proactive: suggests actions Autonomous: plans and executes in a loop
Human Role Initiator and decision-maker Decision-maker with AI assistance Supervisor and approver
Tool Access None (text in, text out) Read access to workspace context Full tool invocation (read + write + execute)
Autonomy Level Zero: follows instructions literally Low: augments human workflow High: operates independently within guardrails
Error Recovery User must re-prompt User reviews and corrects suggestions Agent retries, escalates, or asks for help
Typical Examples ChatGPT (default), support bots GitHub Copilot, SAS Viya Copilot Claude Code, Copilot Coding Agent
Best For Q&A, brainstorming, generation Code completion, editing, analysis Multi-step workflows, research, engineering

 

Table 1: The Autonomy Spectrum — Chatbots, Copilots, and Agents

The Plumbing: MCP, Skills, and Harnesses

If model architecture and autonomy define what an AI system can do, the plumbing defines how it does it. Three concepts have emerged as the foundational building blocks of modern agentic systems, and understanding the distinction between them is essential.

Skills

Skills are reusable, file-based definitions of repeatable workflows. Typically, a skill is simply a Markdown file with natural language instructions. Whereas a tool is a single function (call an API, run a query), a skill can encapsulate an entire workflow: search relevant documents, summarize them, format the output, and present it to the user. If tools are functions, skills are modules composed of multiple functions that an AI agent can understand and follow.

Model Context Protocol (MCP)

MCP is an open standard introduced by Anthropic and has been widely described as “the USB-C of AI” because it standardizes how agents connect to external tools, data sources, and services.

Before MCP, every tool integration was bespoke. Connecting an agent to Slack required custom code. But connecting it to a database required different custom code. Connecting it to a file system required even more custom code. MCP eliminates this fragmentation. Build an MCP server on the tool side once, and every AI client that supports the protocol (Claude, ChatGPT, VS Code, etc.) can connect the same way.

At its core, MCP is a combination of workflows, tools, and reference documentation. However, it lets you to easily connect your AI agent to 3rd-party tooling and APIs without having to know how to call them programmatically. The LLM handles that aspect.

Skills vs. MCP: When to Use Which

Because skills and MCP both extend what an agent can do, it is natural to wonder when to reach for one versus the other.

Use skills when you have a repeatable process or tightly scoped workflow that you want to codify and reuse. “Analyze a clinical trial data set, flag anomalies, and generate a summary report” is a skill. So is “run the quarterly compliance check against these five data sources and produce a formatted findings document.” Skills encapsulate business logic, sequencing, and output formatting into a self-contained unit that can be invoked reliably every time. They are well suited for local execution where security and authentication are not primary concerns: file manipulation, text processing, data transformation, or any workflow that operates on data already within the agent’s environment.

Use MCP when you want to enable users to interact with your tools or APIs via conversational natural language instead of programmatic commands, and you want to do so across different environments. This is the key shift MCP unlocks. Rather than requiring a user to know the right API endpoint, construct a query, or navigate a CLI, they simply describe what they need in plain English.

The agent then translates that intent into the appropriate tool calls through MCP. “Show me last quarter’s revenue by region” becomes a natural language request that MCP routes to the right database query, analytics API, or reporting service behind the scenes. MCP is designed for remote tool calling across authenticated gateways, handling the handshake with external services, managing credentials, and respecting access controls, making it the right choice when the agent needs to reach beyond its local sandbox into enterprise systems, third-party APIs, or cloud services that require governed authentication.

That doesn’t mean the two are mutually exclusive. A skill might orchestrate several MCP tool calls under the hood, or an MCP server might utilize embedded skills. The line is blurry, but generally, skills define the workflow while MCP provides the integration. Table 2 provides more context and differentiation.

Criterion Use Skills Use MCP
Interaction Model Agent executes a codified workflow end-to-end User describes intent in natural language; agent translates to API calls
Data Location Already within the agent’s local environment External systems, cloud services, third-party APIs
Security Posture Low concern (local sandbox execution) Governed authentication across remote gateways
Reuse Pattern Same process, reliably invoked every time Same integration, accessible from any MCP-compatible client
Setup Complexity Create a Markdown file with instructions Configure server, register tools, manage credentials
Composition Skills can orchestrate MCP calls under the hood MCP servers can embed skill-like workflows internally

 

Table 2: Skills vs. MCP — When to Use Which

Harnesses

Another term you may have heard thrown about is “harness”. A harness is the complete runtime infrastructure wrapping an AI model: the orchestration loop, tool dispatch, memory management, context assembly, error handling, and guardrails. Martin Fowler and Birgitta Böckeler of ThoughtWorks crystallized this with a simple formula: Agent = Model + Harness. The model provides the intelligence; the harness provides the structure that makes that intelligence reliable and useful.

A useful way to think about this is with a car analogy. The model is the engine, but the harness is the rest of the vehicle. An engine by itself sits on a shop floor and spins. It takes a chassis, steering, brakes, a transmission, and a dashboard to turn that raw power into something that actually gets you where you need to go safely. The harness provides the steering (context assembly and tool routing), the brakes (guardrails and human-in-the-loop gates), the transmission (memory and state management), and the dashboard (observability and logging). Swap in a more powerful engine and the car gets faster, but without the rest of the vehicle, horsepower is just noise. And crucially, good harness design is modular: as models improve, components of the harness that compensate for model weaknesses can be removed. Table 3 spells out this analogy explicitly.

Car Component Harness Equivalent What It Does
Engine AI Model Provides raw intelligence and reasoning power
Steering Context Assembly + Tool Routing Directs the model’s attention to the right information and actions
Brakes Guardrails + Human-in-the-Loop Gates Prevents unsafe, unauthorized, or costly actions
Transmission Memory + State Management Maintains continuity and context across interactions
Dashboard Observability + Logging Provides visibility into what the agent is doing and why

 

Table 3: The Harness as a Car Analogy

An excellent blog post by the author of bits-bytes-nn, titled “From Prompts to Harnesses: Four Years of AI Agentic Patterns,” traces the evolution of AI engineering rigor over the past few years across three distinct eras:

  •  Prompt Engineering (2022-2024): “What should I say?” The industry believed instruction quality was the sole determinant of AI response success.
  • Context Engineering (2024-2026): “What information should I provide?” The realization that what fills the context window and is available to the AI agent matters more than the prompt itself.
  • Harness Engineering (2026+): “What system should I build?” The acceptance that the design of the entire system is the real problem to be solved. And that intelligent systems can make up for inadequate models.

The core thesis is compelling. As agentic AI proliferates within software development, engineering rigor does not disappear. Rather, it relocates to a higher level of abstraction. This covers writing code, curating context, and designing the environment in which agents operate. For analytics and data science practitioners, this pattern is familiar. It mirrors the evolution from hand-coding statistical models to configuring automated pipelines to designing governed analytical ecosystems. The abstraction level rises, but the need for rigor remains. Table 4 summarizes this recent history.

Dimension Prompt Engineering (2022) Context Engineering (2024) Harness Engineering (2026)
Core Question What should I say? What information should I provide? What system should I build?
Metaphor Writing an email Managing an inbox Designing the email system
Key Metric Response quality (subjective) Retrieval accuracy Task completion rate, cost per task
Failure Mode Blind prompting, non-determinism Context pollution, lost-in-the-middle Orchestration bugs, security incidents
Where Rigor Lives The prompt text itself Context window composition Entire system architecture
Representative Tools ChatGPT, early Copilot Cursor Composer, RAG pipelines Claude Code, Copilot Coding Agent
Required Skills Language sense + domain knowledge Information architecture, retrieval design System design + security engineering

 

Table 4: Three Eras of AI Engineering

Coming Up Next

Autonomy and plumbing determine what an agent can do. The next question is what it knows. Part 3 turns to retrieval: how to ground a model in your own enterprise data, and how standard RAG, knowledge graphs, and the newer PageIndex approach compare when you need answers that are accurate, traceable, and defensible.

Additional Resources




Source link


administrator