AI — the assistant, agents, background tasks, models and MCP
In plain terms: the sidebar group is AI Studio. Overview is its front door — a diagram of what makes an agent, how one gets used, and what this workspace is missing, with at most three next steps. Chat is where you talk to an (ordinary) agent. Agents is the list of the workspace's AI executors — every one either ordinary (chat) or Scheduled (background) and never both, see Chat vs. Scheduled — each with its own page holding every setting it has. Runs is what ran without you. AI settings holds what every agent in the workspace shares: the model providers you have connected your own key for (Models), the outside tools agents may be given (Connections), and what it all cost (Usage).
Renamed and regrouped twice: Models and Connections moved under AI settings 2026-09-09 (/org/:slug/models and /org/:slug/mcp still work — they redirect); Overview was added and "Tasks" became Runs 2026-09-16, after a first-time walkthrough found five sidebar entries named after database tables that said nothing about what connects to what.
Everything here is scoped to one workspace. An agent, a chat, an MCP connection and a provider key all belong to the workspace you were in when you created them, and are invisible from another — with one exception, marked Platform in the agent list: built-in agents the platform provides to every workspace.
The AI runs in its own application (apps/intelligence), not in core-api. It never decides for itself what you may see: it asks core-api "what is this person allowed to do in this workspace" and enforces the answer. That is why the permissions below look like ordinary workspace permissions even though no core-api route checks them — see The permission model.
How to...
Grouped to match the technical sections below — each page's own help icon opens straight into its own group here, not this whole list.
Overview
See how it all fits together — AI Studio → Overview. The section's front door is a diagram, not a list: a model and the connected tools make an agent; chat, a one-off task or a schedule are the three ways to use one; runs are what comes out. Each box says what this workspace currently has, and clicking it opens the page that changes it.
See what is missing — a dashed box is a hole: no tools connected, no agents of your own yet, nothing running on a schedule. That is usually the answer to "why does the assistant only talk in generalities" — it can reach nothing but the platform itself.
Know what to do next — under the diagram, What you can do now turns the same state into at most three actions, most urgent first: a failed run before anything else, then the piece that is missing, then simply asking the assistant something. A step you have no permission for is never shown.
Assistant
Ask something — Assistant in the sidebar → type in the box at the bottom → Enter. The answer streams in as it's written; the arrow button turns into a stop button while it does, and clicking that keeps whatever has already arrived.
Ask how the platform itself works — just ask, in plain words ("what is impersonation," "why can't I see a colleague's payout"). The assistant searches the wiki before answering and links the page it used, rather than guessing from what a model generally believes about software. A question the documentation genuinely doesn't cover gets an answer that says so.
Attach a photo — the paperclip in the message box. Images only, and only in chat — a background task takes text instructions. Whether the answer actually uses the picture depends on the model the current agent runs on.
Change who you are talking to — the agent dropdown at the top right of the chat. It switches mid-conversation: the chat you already have is carried over, the next answer just comes from a different agent. Only ordinary agents are offered — a Scheduled one never appears here (see Chat vs. Scheduled). Nothing is lost and nothing is re-sent.
Start a new chat, or go back to an old one — the chat list on the left. On a phone it is a drawer: the panel icon at the top left of the chat opens it. Each chat keeps its own history; + starts an empty one.
Rename or delete a chat — the chat's own row in that list.
Clear the conversation — the bin icon at the top right. This empties what is on screen for the current chat; deleting the chat itself is the row action above.
Agents
Create an agent — Agents → + (on a phone, the ⋮ menu → Create agent). The very first choice is what kind of agent this is — see Chat vs. Scheduled below, it cannot be changed later. See field guide.
Start from a ready-made agent — Agents → the Ready-made agents tab, beside Your agents. Three that work the team board: a nightly project digest, triage for tasks created since yesterday, and a watcher for tasks nobody has touched in a week. A tab rather than an entry in the + menu (2026-09-16): behind the + nobody could tell what was being offered, and one nobody opens is one that does not exist. Use this one fills in the ordinary create form — instructions, kind and schedule included — and creates nothing until you press Create agent, so you read what the agent will be told before it exists.
Chat vs. Scheduled — chosen once, never changed
Every agent is one of two kinds, picked on the create form and fixed forever after (2026-09-24):
- Ordinary agent — answers in chat. Can never be put on a schedule, and never appears on the Runs page at all, one-off or recurring.
- Scheduled agent — the opposite split, not a superset: it runs on its own, by hand or on a timer, and is never offered in chat — no Talk to it anywhere, on its card or its own page. Its task prompt becomes required at creation (this is the only instruction it will ever get when nobody is watching), and its first schedule is created in the same step — a Scheduled agent can be paused but can never be left with zero schedules, because there would be nothing left that makes it one.
There is no third option and no way to convert one into the other; an ordinary agent that should start running unattended has to be created again as a Scheduled one.
Open an agent — click its card. Everything about one agent lives on its own page: Overview (what it is configured to be, read-only) and Configuration (every setting) — plus, only for a Scheduled agent, a third tab, Runs (what it has done). The tab is part of the address, so a link opens on it.
Change any setting — the agent's page → Configuration. Five blocks, each saved on its own: Identity, Behaviour, Model, Tools, Reference images — and a sixth, Automation (the schedule itself, plus what a run hands back), only for a Scheduled agent: an ordinary agent's kind means there is nothing there to ever configure. On a wide screen the list down the left jumps between them.
Give an agent a picture — the create form, or Configuration → Identity → Upload a picture. Optional: an agent without one wears a colour and a mark derived from itself, so a list of agents never reads as a row of identical robots either way.
Rename an agent, or change its description — Configuration → Identity. Neither line is ever sent to the model; they are how people find the agent.
Change what the agent is told — Configuration → Behaviour. Instructions is what the model reads before every answer, in chat and in background runs alike. Task prompt is the default instruction for a run started without one — ignored in chat, and if you leave it empty every background run has to carry its own.
Change the model — Configuration → Model. Every model is selectable, including one your workspace hasn't connected a key for yet — that is a valid thing to set up ahead of time, not an error; the picker just says — needs a key next to it. Under the picker are the model's own properties — provider, whether it can read images, price per million tokens. Those are the same in every workspace and are set in the platform catalog, not here; provider API keys live in AI settings and never on an agent.
Fine-tune how it answers — Configuration → Model → Generation settings, below the picker. Temperature, Top P and Max output tokens — only the ones the current model actually accepts are shown at all; switch to a model that rejects one and its field simply disappears, saved value and all, rather than sending something the provider would refuse. Left blank, the model's own default is used.
Give an agent access to a tool — Configuration → Tools → tick the connections it may use → Save. Deliberately per agent: connecting an MCP server to the workspace does not hand it to every agent (see MCP Connections). The block underneath lists what every agent can already do without any setup.
Pin reference images — Configuration → Reference images → Add reference image. Images the agent compares every newly attached picture against — a shelf layout, a product's correct packaging. They save the moment you add one. If the chosen model cannot read images the block says so, because nothing pinned would ever reach it.
Choose what a background run hands back — a Scheduled agent's Configuration → Automation → Structured result. Left on Free text, a run's result is whatever the agent wrote; pick a shape instead and the run returns filled-in fields, which is what makes a result something another system can read rather than something a person has to.
Add another schedule, or change one — a Scheduled agent's Configuration → Automation → Add a time, or the ⋮ on a schedule already listed. The toggle on each row pauses that one without deleting it — but the last schedule on a Scheduled agent cannot be deleted, only paused: an agent of this kind can never be left with none.
Delete an agent — the agent's page → ⋮ → Delete, or the card's ⋮ in the list.
Talk to an agent — the card's ⋮ → Talk to it, or Talk to this agent on the agent's page — an ordinary agent only. A Scheduled agent has no such button anywhere: it is not built to answer in a live conversation, by kind, not by mistake.
What a platform agent allows
A card badged Platform is provided by the platform to every workspace — one shared agent, not a copy per organization. It is an ordinary agent by kind — a shared one can never own a schedule (see What cannot be changed, and why) — so its page holds only Overview, which shows what it is set up to be: no Configuration tab and no Runs tab. Nothing about it can be changed from here and it cannot be deleted — the API refuses those writes for a platform agent, since one workspace's edit would land on everybody's. You can still talk to it like any ordinary agent.
It is not unowned, though. A platform administrator configures its name, instructions and model under Platform Administration → AI Assistants (how), and anyone holding platform:ai_assistants:read gets a link to it straight from this page — Configure it for somebody who may change it, Open its settings for somebody who may only look. Every such change is written to the platform audit log, because it lands on every workspace at once.
Which settings exist
Everything an agent has, and nothing it doesn't.
| Setting | Where | What it does |
|---|---|---|
| Kind | Create form only | Ordinary or Scheduled — see Chat vs. Scheduled. Fixed forever. |
| Name, Description | Configuration → Identity | How people recognise the agent. Never sent to the model. |
| Instructions | Configuration → Behaviour | What the model reads before every answer. Required. |
| Task prompt | Configuration → Behaviour | The instruction for a background run. Required for a Scheduled agent; ignored in chat. |
| Model | Configuration → Model | Which model answers. Chosen from the platform catalog, or the workspace's own. |
| Generation settings | Configuration → Model | Temperature, Top P, Max output tokens. Shown only for what the current model accepts. Blank = the model's own default. |
| MCP connections | Configuration → Tools | Which of the workspace's outside connections this agent may use. Nothing by default. |
| Reference images | Configuration → Reference images | Images every new picture is compared against. Needs a model that can read images. |
| Structured result | Configuration → Automation | The shape a background run hands back. Scheduled agents only. |
| Schedules | Configuration → Automation | When it runs by itself — at least one, always. Scheduled agents only. |
Creating an ordinary agent asks for Kind, Name, Description, Instructions and Model — the five it cannot exist without. Creating a Scheduled one asks for the same, plus the Task prompt and its first schedule, right there on the same form: both are made together with the agent, in one step, so a Scheduled agent is never left without either.
Agent runs
Run an agent in the background — Runs in the sidebar → Runs tab → + → pick a Scheduled agent, optionally type an instruction → Run task. It keeps running after you close the page, and the table follows it on its own — Refresh is there if you want it, not because you need it. See field guide.
Find out it finished without watching — you get a notification when a run of yours ends, whether it succeeded or failed. Nothing to switch on; if you'd rather not have them, the group is Agent runs under Account → Notifications.
See what the agent actually did — click the run's row to expand it: the instruction it was given, the steps it went through (a tick, a spinner or a cross per step), and its result.
See what a run cost — the same expanded row, at the bottom. Only visible if you hold org:billing:read; a run that finished before any usage was recorded says so rather than showing 0.
Repeat a run automatically — Schedules tab → +. Either every N minutes or a cron expression, with the timezone it should be read in. The tab itself only exists for people who can manage agents — see field guide.
When a schedule says "Cannot run" — its agent's model has no key connected, or was switched off. It is never refused outright: a Scheduled agent on such a model is still created — with its very first schedule already disabled, and a toast says so plainly at the moment of creation, so the agent exists ready to switch on the moment a key is connected. What a broken schedule cannot do is be switched on, and a run cannot be started by hand either: the screen says what is missing at the moment you press the button. An already-stuck schedule shows the reason on its own card, with the place that fixes it — Connect a key opens AI settings, Change the model opens the agent. Switching it off, or changing its timing, is always allowed. An OpenRouter model needs no key of your own; if the installation has none, an administrator sets it under Platform Administration → AI models.
Pause or remove a schedule — the switch on the schedule's own card flips it between Enabled and Disabled without opening anything; ⋮ → Delete removes it for good. Either way, runs that already happened stay in the Runs tab.
Run a task — field guide
- Agent — only Scheduled agents are listed (see Chat vs. Scheduled). If the list is empty, that is what to fix first — create one, or turn an existing job into one, on the Agents page.
- Instruction (optional) — leave it blank and the agent's own Task prompt is used, which for a Scheduled agent always exists.
Create/edit a schedule — field guide
- Agent — a Scheduled agent already has its first schedule from the moment it was created; this dialog is for adding a second one, or editing an existing one.
- Repeat — Every N minutes for a simple interval, or Cron expression for anything calendar-shaped (standard 5-field syntax, e.g.
0 9 * * 1for every Monday at 09:00). - Timezone — a cron expression means nothing without one; this is the zone the expression is read in, not your browser's.
- Remember previous runs — off by default, and worth leaving off unless the job actually needs it. On, every run of this schedule shares one conversation, so a daily digest can know what it already reported instead of starting from nothing each morning. The cost is that the conversation keeps growing, and so does what each run costs. It belongs to the schedule, not the agent: two schedules on the same agent remember separately.
- The Enabled switch is on the schedule's card, not in this dialog. A disabled schedule is kept, it just stops producing runs.
Model Providers
Use your own API key — AI settings in the sidebar → Models tab → Connect on Anthropic, OpenAI, Google or OpenRouter → paste the key → Connect. The dialog links out to where each provider issues one and shows the shape that provider's keys have. Then open an agent and pick one of that provider's models: the key alone changes nothing, since an agent runs whatever model its own page says. Everything your agents run on that provider is then billed to your own account instead of going through the platform's shared gateway. OpenRouter is the one exception worth knowing: the platform already reaches it with a shared key on everyone's behalf, so connecting your own there only replaces whose bill it lands on — nothing stops answering if you skip this one.
Add a model of your own — once a provider is connected, + Add model on its card → pick the exact model id → Add. This is your row: it appears in every agent's model picker for this workspace only, billed on your own key, and only you can edit or remove it — see Your own models further down the page. The platform's shared catalog (Platform Administration's own list) is unaffected either way.
Stop using your own key — Disconnect. Agents pinned to a model from that provider stop working until you reconnect — they do not silently fall back to another model. Any custom model you added on that provider's key goes with it.
MCP Connections
Connect a ready-made connector — AI settings in the sidebar → Tools tab → Predefined connectors → Connect on the one you want (Google Workspace today). This sends you to that service's own authorization page and back; you never paste a key.
Connect your own MCP server — + (on a phone, the ⋮ menu → Add connection). See field guide.
Turn a connection off without deleting it — the card's ⋮ menu → Disable. The card stays, greyed out, and its tools disappear from every agent until Enable.
Let an agent use a connection — this page alone does not. Go to Agents, open the agent, and tick the connection on its Tools tab.
Add/edit an MCP connection — field guide
- Name — yours, shown on the card.
- Server URL — the MCP server's endpoint.
- Authentication — No auth, or Header auth with a Header name and Header value. The value is stored encrypted and never shown again: when editing, leaving it blank keeps the current one.
- Enabled — same meaning as the Disable action above.
Spend
See what the AI cost — Usage in the sidebar. Pick a period and whether to group by model, agent or day.
Read the caveat on the total — a model with no price set contributes tokens but no cost, and its row says no price set rather than $0.00. The total then carries a note saying so: it is a floor, not the full bill. Prices are set per model in the platform's own catalog.
The permission model
| Permission | What it opens |
|---|---|
ai:chat:use | Assistant, and one-off runs in Runs |
ai:agents:manage | Agents and Model Providers; also the Schedules tab |
ai:mcp:manage | MCP Connections |
org:billing:read | Usage, and the cost line inside an expanded run |
AI settings is one sidebar entry over two permissions: it appears for either ai:agents:manage or ai:mcp:manage, opens on the tab that permission covers, and shows only the tabs actually held. Someone with just ai:mcp:manage therefore sees Tools and no Models tab — merging the two entries did not merge the two gates.
In the seeded roles, OWNER and ADMIN hold all four; MEMBER holds ai:chat:use only — so a member can chat and start a single run on a Scheduled agent someone else already created, but cannot create an agent of either kind, connect tools, commit the workspace to a recurring spend, or see what any of it cost. The one deliberate asymmetry inside that set: starting a single run on an existing Scheduled agent is ai:chat:use like chat itself, while creating one, or adding a schedule to it — an open-ended commitment to spend — is ai:agents:manage.
These four are the only permissions in the whole catalog that no core-api route checks. apps/intelligence resolves them through GET /organizations/:id/access and enforces them itself, which is also why they are invisible to core-api's own guards. See docs/ai/ADR-001-ai-platform-architecture.md Decision 3, and docs/iam/PERMISSIONS_CATALOG.md.
Assistant (/org/[slug]/assistant)
The chat page: a chat list on the left (a drawer below md), the conversation in the middle, the message box at the bottom, and the current agent in a dropdown at the top right. The header reads Platform Assistant whichever agent is selected — the dropdown next to it is what says who is actually answering. (Everything on this page was hardcoded English until 2026-09-06; it is fully translated now, header, empty-state prompts, thread list and streaming errors included.)
Message history is not stored by hand: Mastra's own Memory (Postgres-backed, in the separate postgres-intelligence database) owns threads and messages. The platform's own ThreadMeta row holds only what Memory is agnostic to — which workspace and user a chat belongs to, its title, and which agent is currently assigned. That last field is what makes switching agent mid-thread a non-event: the same thread id is simply handed to a different agent.
An attached photo is uploaded to core-api first (POST /organizations/:id/chat-attachments, the platform's ordinary file storage) and then referenced by id in the message — the picture itself never travels through the chat request.
Before answering a question about the platform itself, the assistant searches the wiki (search-help/read-help-section, both MCP servers) and cites the page it used, rather than answering from a model's general beliefs about software it has never seen.
Reached by ai:chat:use.
Agents (/org/[slug]/agents)
An agent is a configurable AI executor, not a chatbot with a personality: a set of settings somebody chose, which the platform then runs.
Every agent is one of two kinds, chosen once at creation and never changed afterward (Agent.kind, 2026-09-24) — see Chat vs. Scheduled for the how-to. An ordinary agent only ever answers in chat. A Scheduled agent only ever runs unattended — by hand, on a timer, or both — and is never offered in chat at all. The two populations do not overlap, and nothing converts one into the other: this is a real, stored column, not an inferred state, and it is what every other kind-dependent rule in this section ultimately rests on — the Runs tab, the Automation block, the chat agent dropdown, the "Talk to it" button.
The list has two tabs — Your agents and Ready-made agents (three built-in ones prefilled onto the create form, see How to...). One card per agent, carrying the same facts in the same order so two side by side still compare: the model it runs on, and — for a Scheduled one — a clock glyph and how its last run ended. Which MCP connections each agent may use is deliberately absent from the card — that is only readable one agent at a time, and one request per card to fill a single line is not worth it.
An agent is six things that matter to the model — instructions, a model, its generation settings, its ticked MCP connections, any pinned reference images, and — Scheduled only — a task prompt plus at least one schedule and an optional structured result shape — and two things that matter only to humans, its name and description. The description is deliberately never sent to the model.
Reached by ai:agents:manage. Delete takes the agent out of the list and, for an ordinary agent, out of every chat's agent dropdown, with no undo in the interface (the row itself is soft-deleted, so past runs that reference it stay readable). An open chat that was using a deleted agent keeps its history — pick another agent from the dropdown to carry on.
What an agent can reach
Every agent can look at the platform itself — who you are, which workspaces you are in, who is in them, what roles and subscriptions they have — and can search this wiki for how the platform actually behaves, citing the page it used. No setup, no MCP connection: those tools are always there.
If you are part of the P4P team, an agent can also work with the team's own things — departments, projects, tasks and their statuses, the activity feed, the HR pool — and, separately, read your calendar. Which of those it gets is decided by your access, not by the agent: an agent sees a tool only if you could use it yourself, and creating or changing a task needs the same permission it would need from you. Someone who cannot open the team board has an agent that cannot either, and it does not fail at that — the tools are simply not there.
Two of those tools exist for agents that work the board rather than answer questions about it. One asks for the board as a summary — per project, what was created or touched recently, what is overdue, what nobody has picked up, what has not moved in a week — instead of pulling every task in full, which is what makes a nightly digest cheap enough to run every night. The other posts in a project's discussion, so what the agent wrote lands where the people it concerns already are, with the usual notification, rather than in a run log nobody opens.
The same rule holds when nobody is watching. A scheduled run works with the access of whoever set the schedule up, so it can reach exactly what that person can and stops working if they lose it. Every one of these calls is recorded, so "what did an agent do on my behalf" has an answer.
What an agent gets back is deliberately narrower than what you see on screen: names and roles rather than the team's email addresses, a person's own account rather than their phone number and address. Anything the assistant told a model stays in that conversation, and there is no taking it back — so the answer is only ever what the question needed.
Agent runs (/org/[slug]/agent-tasks)
Two tabs. Runs is the table of everything that has run — agent, status (Pending, Running, Succeeded, Failed), trigger (Manual or Scheduled), created and finished times — with each row expanding into the instruction it was given, its steps, its result, and, for org:billing:read holders, its token usage and cost. Schedules is the recurring side, and only appears for ai:agents:manage holders. Both tabs offer only Scheduled agents — see Chat vs. Scheduled; an ordinary agent cannot appear on this page at all, in either tab.
The whole page exists because a Scheduled agent only works while nobody is watching it. A run keeps going after the page is closed; the table is the record of what happened while you weren't there.
The list keeps itself current, and how often depends on whether anything is happening: every few seconds while a run is in flight, and much more slowly when the table is idle — but never zero, because a schedule can start a run while you have the page open and nothing else would tell it. When a poll finds nothing new the table is left exactly as it was, so it will not move under you while you read. You also don't have to keep the page open at all: a finished run notifies whoever it ran for.
A background run works with the same access as the person it runs for — the one who started it, or, for a schedule, whoever set the schedule up. It cannot see anything that person couldn't, and if they lose access to the workspace the run stops rather than continuing under their name.
Model Providers (/org/[slug]/ai-settings/models)
Four cards — Anthropic (Claude), OpenAI (ChatGPT), Google (Gemini), OpenRouter — each Connected or Not connected.
By default the platform reaches every model through one gateway (OpenRouter) with its own key, and a workspace needs nothing here at all. Connecting your own key changes only how the same models are reached for this workspace: directly, on your billing account. The key is encrypted at rest and never shown back — disconnect and reconnect to replace it. OpenRouter is the one card with a platform-wide fallback behind it; the other three have none — an agent on Claude with no Anthropic key connected simply cannot run, there being no installation-wide Claude key to fall back to.
That gateway key belongs to the installation, not to any workspace — it is set once per deployment, from Platform Administration → AI models. An installation with no key set answers nothing anywhere, and the provider returns an empty completion rather than an error, so the assistant simply goes quiet. The page says so outright when that is the case, and connecting a key of your own is a way around it that needs nobody's help.
Connecting a key switches nothing by itself. Every agent runs the model chosen on its own page; a new key only makes that provider's models selectable there. The Models linked to agents card at the top of this page names the model each agent is on, which provider it runs on and whose key pays for it — the platform's or yours — with a clock glyph beside a Scheduled agent's name (it will never appear in chat, so that's the one fact worth a tooltip here).
Once a provider is connected, add your own model on top of it (2026-09-24) — the + Add model button on that provider's card, listed under Your own models below. This row is yours alone: it appears in every agent's model picker in this workspace only, is billed to your own key by construction (the platform never has a key it could run on), and only this workspace can edit or delete it. The platform's own shared catalog — every model an administrator has curated in Platform Administration — is unaffected and stays available alongside it. Disconnecting the underlying key removes any custom model built on it too.
Reached by ai:agents:manage.
MCP Connections (/org/[slug]/mcp)
MCP (Model Context Protocol) is how an agent gets tools — read a mailbox, fetch a document, post a message — instead of only producing text. This page has two halves:
- Predefined connectors — maintained by the platform and authorized over OAuth, so there is no URL or secret to type. Google Workspace (Gmail and Drive) today.
- Your own connections — any MCP server you can reach: a name, a URL, and either no authentication or a single header whose value is stored encrypted.
Connecting a server here does not give any agent access to it. That is the one rule on this page worth remembering: access is granted per agent, on the Agents page, by ticking the connection in the agent's own MCP connections list. Deny-by-default, deliberately — an organization-wide connection would otherwise silently widen what every existing agent can reach the moment it is added.
Reached by ai:mcp:manage. Disable removes a connection's tools from every agent at once while keeping its configuration; Delete takes it away for good as far as the interface is concerned (the row is soft-deleted, and agents that had it ticked simply stop seeing its tools).
The reverse direction — external AI clients such as ChatGPT or Claude Desktop reaching into P4P — is not this page. That lives in your account, under Connected apps.
Testing this module
Every page in this section has its own unit-test file (pages/org/[slug]/{assistant,ai/index, agents/index,agents/[id],agent-tasks/index,ai-settings/{models,mcp,usage}}.test.ts).
scripts/e2e/ai-studio/ (overview.mjs, agents.mjs, settings.mjs, runs.mjs, platform-models.mjs, blocked-schedules.mjs, chat-model-fix.mjs) is a standing, re-runnable e2e suite driving every control in this section against a real dev server, against its own throwaway workspace (ai-studio-e2e). scripts/e2e/agent-runs/ covers a background run's identity and a schedule's completion notification through a real model call; scripts/e2e/assistant/guide.mjs asks the real chat real questions and judges the tool trail, not just the prose, including whether it actually cites the wiki. scripts/e2e/mcp/ covers both MCP servers (internal and external) over real HTTP. None of this existed before 2026-08-12.
A real OPENROUTER_API_KEY works locally, but two environment gotchas both cost a debugging session the first time: OpenRouter is unreachable directly from some networks (a proxy plus NODE_USE_ENV_PROXY=1 on Node ≥ 24 fixes it — see apps/intelligence/CLAUDE.md), and the platform's own free default model is rate-limited per hour on the shared pool, so a 429 there is not a broken key.
Found and fixed while writing this section (2026-09-25). Two real gaps, not test flakiness:
- Saving any edit on Platform Administration → AI models answered 503 — "Could not record this change in the audit log, so it was not applied." Decision 20 added
supportsTemperature/supportsTopP/supportsMaxOutputTokensto what an edit can change, but core-api's own allowlist for what the audit trail may record (AI_MODEL_FIELDS) was never told about the three new fields, so the very form that sends them tripped its own guard on every single save. Fixed by adding them to that list (apps/core-api/.../record-ai-model-change.dto.ts); confirmed live and pinned byplatform-models.mjs. runs.mjsandblocked-schedules.mjspredatedAgent.kindand built their fixtures the old way — an ordinary agent, orallowAutomatedInvocationset by hand — which can no longer carry a schedule at all. Rewritten onto a real Scheduled fixture (_shared.mjs'sensureScheduledAgent) and Decision 21's actual create-time behaviour (a blocked Scheduled agent is created anyway, paused — only a second schedule on top of it is still refused outright).