MOD · Dispatch · Blog S/N · 8 MIN · JUN 4, 2026

8 min read

A role is not a prompt

Most AI agents are one paragraph of stage direction. Team-X defines an agent as a schema-validated role spec, and the runtime enforces some of it. This is a field-by-field audit of which parts, at v3.2.1.

  • role-specs
  • least-privilege
  • agent-security
  • architecture

Most AI agents are a paragraph of stage direction. “You are a senior engineer.” The model is asked to behave, and if it does not, nothing stops it.

I built Team-X on a different bet, and this post tests whether the code backs it up.

Team-X, an open-source, local-first desktop app for running AI-agent organizations, defines an agent as a role spec: a markdown file with schema-validated YAML frontmatter and a prompt body. Some frontmatter fields are enforced by code, such as level and tool lists. Others only describe. I audited every field at tag v3.2.1 and sorted them.

What does a real role spec look like?

It looks like a config file with a prompt attached. There are 57 files under the roles directory at v3.2.1: 55 across six hierarchy levels and two system roles that are never hired through the UI.

Here is the head of the CEO spec, copied verbatim:

id: chief-executive-officer
name: Chief Executive Officer
level: officer
reports_to: [board]
manages: [coo, cto, cmo, cfo, cpo]
preferred_model_tier: high
preferred_providers: [anthropic, openai, ollama]
fallback_providers: [groq, openrouter]
preferred_context_window: 200000
tools_allowed: [browse, email, calendar, context7, episodic-memory]
tools_denied: [shell, filesystem_write]
decision_authority: final
escalates_to: []
kpis: [revenue, team_health, product_vision, runway, customer_love]

Everything after the closing delimiter is the system prompt. The parser validates the frontmatter with a zod schema. Required fields include id, name, level, preferred_model_tier, decision_authority, and temperature; the list fields default to empty. An unknown level or a missing field throws before the role loads.

The pack is also signed. The loader verifies an Ed25519 signature over the pack directory and runs in strict mode in packaged builds, warn mode in dev. If the schema is where authority lives, the schema file needs tamper protection.

Which fields does the code actually enforce?

Three groups: level, the tool lists, and nothing else that grants or denies anything. The rest of the frontmatter either scores, describes, or is not read. Here is the full sort, from a grep of apps and packages at v3.2.1 for every field name:

FieldWhat reads it at v3.2.1Status
levelThe hire tool gate, the project-decomposition approval check, the manager-inversion guardEnforced, with a spelling gap (below)
tools_allowed, tools_deniedAuthority resolver, tool pre-filter, MCP host call checkEnforced for MCP tools only
capabilitiesPlanner scoring of which employee fits a subtaskPicks an assignee; grants nothing
decision_authority, escalates_to, kpisSchema validation; no other reader foundDescriptive
reports_to, managesSchema validation; reporting lines are org edges set elsewhereDescriptive
preferred_model_tier, preferred_providers, fallback_providersProvider factory header says it does not consult themNot read from the role
BodyRendered with variable substitution into the system promptPrompt only

Read the right column twice. decision_authority: final on the CEO is a label. The provider factory states that it does not consult the model preferences, and I found no other code that passes the role’s values into routing. If you read decision_authority as a permission, the code does not agree with you.

The CEO spec shows a second honesty problem. manages lists coo and cpo, but the officer role ids are chief-operating-officer and so on, and git grep finds no role with the id cpo or board. Dangling references are harmless only because nothing reads the field. That is not a virtue. It is a field I should either enforce or delete.

What happens when an agent tries a denied tool?

It never sees the tool, and if it names the tool anyway, the host refuses. Two checks, both in the main process.

First, when the orchestrator builds an agent’s tool set, it filters MCP tools by the resolved lists. A denied name is dropped. A non-empty allowlist drops everything not on it. The code comment says why: so the model never sees denied tools.

Second, the MCP host checks again at call time. Its comment calls it the critical trust-boundary check, and says agents cannot bypass tool restrictions via prompting. A denied call returns a denied status and emits an authority.violation event on the bus.

The Model Context Protocol specification says the protocol “does not mandate any specific user interaction model.” The protocol leaves tool policy to the application that hosts it, so the host is where Team-X puts the rule.

Authority belongs in a schema the runtime can enforce, not in prose the model is asked to respect.

The resolver that produces those lists is layered. Role defaults sit at precedence 0, then extension, company, and employee grants, and platform hard-denies at precedence 4. So a role’s tools_denied is the lowest layer. A company or employee grant outranks it, and only a hard-deny outranks those. At v3.2.1 the composition root builds the resolver with no hard-deny lists, so that top layer is wired but empty.

Two more limits. The built-in orchestrator tools, which let an agent message a colleague or list colleagues, are not subject to the lists. They filter MCP tools only. And the lists match on the MCP tool’s name. The CEO’s entries read like server names, not tool names, so whether they match anything depends on which servers an operator installs and what those servers call their tools.

Is that least privilege?

The mechanism is. The shipped defaults mostly are not.

Saltzer and Schroeder stated the principle in 1975: “Every program and every user of the system should operate using the least set of privileges necessary to complete the job.” NIST’s glossary gives the modern form: each entity is granted “the minimum system resources and authorizations that the entity needs to perform its function.”

For agents, the failure has a name. The OWASP Top 10 for LLM Applications 2025 defines Excessive Agency as “the vulnerability that enables damaging actions to be performed in response to unexpected, ambiguous or manipulated outputs from an LLM, regardless of what is causing the LLM to malfunction.” Its listed root causes are excessive functionality, excessive permissions, and excessive autonomy. Its prevention advice is blunt: “Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not.”

That sentence is the whole argument for a role spec with enforced fields. A persona paragraph relies on the model to decide. A denied-tool list does not ask it.

The cost of this design is real. Someone has to write the lists, and someone has to keep them current as MCP servers change. A prompt is free to edit and free to ignore. A schema is neither, and that is the point.

Anthropic’s engineering team draws the same line between a workflow and an agent: agents are “systems where LLMs dynamically direct their own processes and tool usage”. When the model directs tool usage, the boundary around that direction has to live outside the model.

Spec quality matters at the system level too. A study of multi-agent failures, Why Do Multi-Agent LLM Systems Fail?, builds a taxonomy of 14 failure modes in three categories: system design issues, inter-agent misalignment, and task verification. I am not claiming a number from it. I am noting that design-time decisions get their own category.

Now the part that does not flatter me.

At v3.2.1, 56 of the 57 role files set both tools_allowed and tools_denied to empty arrays. Only ceo.md sets either. And an empty allowlist means unrestricted: the host’s rule is that the tool must be on the list only if the list is non-empty. That is the opposite of a least-privilege default.

So the honest reading is this. The enforcement is real. The official pack barely uses it.

What else is wrong at this ref?

Two things I found while writing this.

First, the level gates are partly decorative. The registry of planning tools is built from lists that include every level, with the comment “All employees can plan”. The real gate is a runtime check that refuses decomposition when the actor’s level ranks below the configured approval level, which defaults to management.

Second, the spelling gap. The role files and the schema enum spell the VP level senior_management. The hire gate and the approval ranking spell it senior-management. The level normalizer lowercases, trims, and turns whitespace into hyphens. It does not touch underscores. Seeding stores the frontmatter value as written.

Traced through the source, a VP is refused at both gates. The hire gate is a plain set lookup, HIRE_LEVELS.has(args.actorLevel), and senior_management is not in the set. The approval check reads the level from a rank table with rank[level] ?? -1, finds no such key, and falls back to a rank below every real level.

That is a finding from reading the code. I have not reproduced it in a running build, and the first thing it needs is a test that hires a VP from the shipped pack and asks it to hire someone.

What should you do with this?

Treat any agent framework’s role definition as two documents: what the model is told, and what the runtime enforces. Then check which fields land in which.

  1. Open a role file and mark each frontmatter field as enforced, scored, or descriptive, as the table above does.
  2. Write a tools_allowed list for any role that touches the outside world. An empty list is unrestricted.
  3. Test a denial. Add a tool to tools_denied, ask the agent to use it, and look for the denied result and the authority.violation event.
  4. Do not read decision_authority or kpis as controls. At v3.2.1 they are labels.
  5. In any other framework, ask where authorization runs when the model asks for a tool. If the answer is the prompt, it is a suggestion.

The role catalog is browsable at /roles/. The code is in the Team-X repository, and I would rather you find the next gap than have me hide it.

Frequently asked questions

What is a role spec in Team-X?

In Team-X, an open-source local-first desktop app for running AI-agent organizations, a role spec is a markdown file. Its YAML frontmatter is validated by a schema, and its body becomes the system prompt. The frontmatter holds level, tool allow and deny lists, capabilities, and several descriptive fields.

Does Team-X enforce tool restrictions in code or in the prompt?

Team-X enforces the tool allow and deny lists in the main process, not in the prompt. Denied MCP tools are filtered out before the model sees them, and the MCP host refuses a denied call again at call time. Built-in orchestrator tools are exempt from these lists.

Which role spec fields does Team-X actually read at runtime?

At v3.2.1, Team-X reads level for hire, planning, and manager checks, the tool allow and deny lists for MCP tool access, capabilities for assignee scoring, and the body as the system prompt. I found no runtime reader for decision authority, KPIs, escalation targets, reporting fields, or the preferred provider fields.

A persona is a suggestion. A denied tool is a fact.
AUTHOR Rocky Elsalaymeh PUBLISHED Jun 4, 2026 ● LIVE