# Mindset AI — Full Site Content > This file contains the complete text content of https://www.mindset.ai in markdown format, intended for LLM consumption. For a summary, see [llms.txt](https://www.mindset.ai/llms.txt). --- --- ## Innovate fast with AI agents in safety-critical products, with confidence. *For mid-market to Enterprise B2B SaaS product & engineering teams* In a regulated product, every agent you build for customers means more compliance to navigate, and innovation stalls. Mindset is the agent harness that handles it: build, run and improve user-facing agents with all AI-first compliance capabilities built in. ### Built for regulated training & compliance sectors EHS & Safety SaaS Platforms, Construction & Contractor Safety Training, Oil, Gas & Energy Training, Healthcare Compliance Training, General Compliance & Professional Qualifications ### In a regulated product, every agent you ship is a new compliance problem to solve. #### A review for every agent or new innovation that comes out Your skills, MCPs and automations run on individual machines, so no one has a shared view of what exists and no guardrails exist. #### Hard to prove it's safe Thumbs up/down and agent's logs aren't enough. You can't show what it did, on what data, who updated what, or that it won't give a customer wrong advice. #### A re-review for every change A prompt tweak, a content update or a model swap re-opens compliance, so improving an agent is as slow as building it, and it usually waits in the backlog. ### Mindset is your agent harness that lets you innovate with user-facing AI fast, with all compliance capabilities built in. Certified against: ISO 27001, GDPR, EU AI Act, Audit trail, Versioning & rollback. ### How Mindset makes it easy to innovate with agents inside safety-critical, regulated products #### Certify the platform once, ship every agent without its own review Identity, audit, isolation and data controls live across the agent harness against ISO 27001, PII protection, GDPR and the EU AI Act's high-risk obligations. Every new agent inherits it all, so you pass every test. #### Mindset lets you build & improve agents visually, so you unlock your entire team for building & maintaining safely Assemble and adjust agents in a visual builder, with engineering's guardrails enforced, so product, support and ops build and maintain the fleet too, not just your core engineers. Every agent is composed from same building blocks, so they build safely. #### Mindset lets you run hundreds of test-cases & monitors agent performance so you have confidence its giving the right answers Mindset monitors real usage across every agent and customer and shows you what's working and what isn't — so you improve from the data, not guesswork, and change it in conversation, not a ticket. #### Versioning and rollback, so any change is reversible, monitored and reportable Every agent, config and rule is versioned: who changed what, from what to what, and when. Roll a change back instantly, or fork a signed-off agent into a new version without redoing its review. #### An independent audit trail that proves what every agent did The platform records every action and change. Click any object, a connection or a single call, through to who ran it, when, on what data and the conversation it belonged to, then download it. #### Mindset keeps your independent across all ecosystems, lets you self-host & gives you a perpetual license Any model, any ecosystem, built on open protocols, self-hostable. No eco-system lock-in. and complete control #### Where does our data go? Self-host on your own infrastructure or run in a chosen region (US, UK, EU), and control what data reaches a model. Your data stays in your environment by default — nothing leaves your infra unless you allow it. #### Which standards does the platform meet? The platform carries ISO 27001, GDPR and the EU AI Act’s high-risk obligations (automatic logging, human oversight, record-keeping), with an independent audit trail and versioning underneath. #### Do we have to re-certify every agent we ship? No. Identity, audit, isolation and data controls live in the platform, so a security review clears the platform once and every agent you build inherits it — instead of each agent going through its own review. #### Can we bring our own model and keep control? Yes. Any model, any ecosystem, built on open protocols and self-hostable. No lock-in, and you keep a perpetual licence to what you build. #### How do we prove what an agent did? The platform records every action and change. Drill from any object, connection or single call through to who ran it, when, on what data and the conversation it belonged to — then download it. --- ## The delivery layer for AI and automation agencies. *For agencies of 10 to 200 people* Turn the skills, MCPs and integrations you write for clients into automations that run inside every client's business, and manage, improve and optimize them from one place. ### Skills and MCPs made it easy to automate your client's team, and hard to manage it once you hand it over to them. You can build it in days. Once it is with them you cannot see what they are running, update it, or tell whether it still works. #### Reuse what you build across clients, so each one starts faster and you can demo it live Take the build from one client, customise it for the next, and install it with their own credentials, permissions and data. #### Build functions for the steps that must not vary, so every run is reliable An LLM running a skill or an MCP is never deterministic. #### Ship with enterprise auditing, compliance and data controls, so you can win bigger clients What you deliver arrives with auditing, compliance and data controls already in it. #### See the health of every automation you built on one board, so you fix the break before the client does Every run, failure and slow response across every client sits on one board, with replay so you are not guessing from a screenshot. #### Attribute every model call, so you can bring their bill down and show the difference Spend is attributed by client, agent, user and model as it happens. #### Publish it into ChatGPT, Claude, Slack and Teams, so nobody changes how they work What you build is published as an MCP server, so it turns up inside ChatGPT, Claude, Slack or Teams, wherever their team already works. ### Bring one client workflow & start today ## Code-only approaches are great for a few agents. Not a fleet. *For mid-market to enterprise B2B SaaS product & engineering teams* Code-only approaches for building agents can't keep up as an agent fleet grows: slow to ship more, unsure what's working, unable to improve without a ticket. Mindset is the harness for agents your customers use - build from composable blocks, improving from real usage, fixes you approve in conversation, not tickets. ## Code-only makes building and managing a fleet of agents really hard. ### One queue, endless demand Sales, CS, product and a big customer all want one, but every build funnels through the same engineers. The business always wants more, faster. ### Is it even working? It's live, but thumbs up/down and evals tell you nothing. You can't see whether it's actually delivering value, or what to change to make it better. ### A ticket for every fix A new MCP or prompt tweak is a code change, so every fix waits in the backlog and takes valuable engineering time. ## The agent harness that takes you from a few agents to a whole fleet anyone can manage - built, run and improved from one simpler place. And it improves itself, while you sleep - a closed loop that turns real usage into safe, approved improvements. ### Improves itself, while you sleep - **Spots failures**: Reads every chat, flags what's failing. - **Writes a test**: Turns a failure into a test you approve. - **Suggests a fix**: Drafts prompts & tools, tested per model. - **Checks it's safe**: Won't ship anything that breaks what works. - **You approve**: One click to ship, instant rollback. Repeats continuously, learning across every customer. ## How Mindset takes your code-only infra from a few agents to a fleet your entire team can manage ### Instead of a ticket, engineer & prioritisation for everything, Mindset lets your entire team build & improve agents visually. Assemble and adjust agents in a visual builder, with engineering's guardrails enforced, so product, support and ops build and maintain the fleet too - not just your core engineers. ### A code-only approach makes you build every agent from scratch. Mindset assembles them from building blocks. Memory, RAG, tool-calling, orchestration, generative UI - every agent is composed from the same building blocks, so each one you ship makes the next faster. ### A code-only approach can't tell you if the agent's working without stitching together multiple systems. Mindset shows you exactly what to fix. Mindset monitors real usage across every agent and customer and shows you what's working and what isn't - so you improve from the data, not guesswork, and change it in conversation, not a ticket. ### Code-only frameworks leave generative UI to you, built agent by agent. Mindset gives you the tools to scale it across the fleet. A framework gives you low-level primitives. Mindset lets you build generative UI widgets straight into your agents and tools - via Widget Studio or MCP - and render them anywhere (chat, your UI, Microsoft Teams), with shared state so each agent sees and controls the page. ### A code-only agent framework locks you into an ecosystem. Mindset keeps you independent and gives you a perpetual licence to the code. Built on open protocols and exported under a perpetual source licence, so the IP stays yours and adds to your valuation. ### A code-only agent framework gives you docs, Discord and complexity. Mindset gives you experts who guarantee delivery. Our engineers help you build the first agents alongside your team to an agreed deadline and hand over the patterns - so your people become the AI experts. ## Trusted by technology & product leaders of leading B2B SaaS companies > "Everyone's shipping 'AI agents' right now. Most of them are a chat box with a fresh logo. Last week at the SHRM conference I launched ClearCo's Agent Platform - Agent Plans, multi-step workflows that run with 100% repeatability. That's the line between a neat demo and real automation." > — Arnaud Grunwald, Chief Product Officer, ClearCompany > "We needed a toolkit that gave us the flexibility to deploy and integrate in different ways across web, mobile, and new products. It had to be something our product and content teams could run directly after implementation, not something that created a permanent engineering dependency. We built the agent experiences we needed, our non-technical teams run the day-to-day, and engineering supported with implementation so the core platform roadmap was not sacrificed." > — Michaela Heigl, CTO, Mindtools Kineo > "Developer time is our most limited asset. If we were going to do this across our 20 products, we needed our product managers to actually own it too, within a secure and scalable architecture." > — Kevin O'Hara, EVP Platform, HealthStream > "Neither of our products had an agent layer, and my team was flat out on a big project through the end of the year. We needed something that could give each program its own agent with access to the right content, across 2,000 programs, without any data leaking between clients, and that could sit inside both our platforms without us having to rebuild them. I also wasn't willing to add five vendors to get there. I wanted one partner who covered the lot." > — Emma Thompson, CTO, Hult Ashridge ## Build, deploy and improve - from one place. ### Build: Build agents and gen-UI experiences from building blocks, in a visual studio. - Set the agent's purpose, personality and guardrails, choose or add your own LLM, and give it Plans for multi-step orchestration and Skills for reusable expertise. Ingest content via our APIs, connect your own RAG, or sync from a source like SharePoint. - Connect your systems: expose your own APIs as MCP servers in your infrastructure, or bring any third-party MCP server. Data never leaves your infra. - Build generative UI widgets in the Widget Studio that render inside the chat - cards, forms, comparisons and charts - styled to your own design system. ### Deploy: Have complete control over how you deploy inside your product, with APIs and SDKs. - Serve thousands of customers from one agent: each user's session is assembled at runtime via an API call to Mindset, binding that customer's own knowledge, tools, skills and permissions, so one agent runs as thousands of isolated versions. - Integrate the Agent Chat SDK for a ChatGPT- or Claude-grade experience out of the box - voice, multimedia, canvas - or go headless and render agents in your own front-end components. - Let agents see and act on the page, not just talk: pass your page state so agents see what the user's looking at, register front-end actions so agents control the existing UI, and have them start conversations off your own backend events. ### Improve: Know whether agents are actually working, and fix them without a ticket. - See exactly how agents are interacting with users - dashboards capture every conversation: what's asked, sentiment, themes, content gaps and where they fall short. - Anyone - product, support, ops, engineering - can improve an agent without a ticket: change a prompt, guardrail or content in the studio and test the change before it goes live. - Surface the data anywhere you work - pipe reports and events to your own data lake so your customers can view it, or use Mindset's built-in dashboards and reports. ## Frequently asked questions ### Does Mindset replace my existing stack or framework? No - it plugs into what you already have (bring your own framework, RAG, MCP servers) and adds the assemble/deploy/improve layer on top. ### Do we keep ownership of what we build? Yes - open protocols, widgets export as standard React, fork the framework under a perpetual source licence. The IP stays yours. ### Where does our data live? In your infrastructure. Your MCP servers run in your environment; agents only work on scoped data; raw records never leave. ISO 27001, GDPR-ready, EU AI Act-aligned. ### Can non-engineers really build and change agents? Yes - engineering wires systems in once and sets guardrails; product, support and ops then build and improve in the visual studio, no code, no deploy. ### How do we serve lots of customers from one agent? Each user's session is assembled at runtime via a single API call, binding that customer's own knowledge, tools, skills and permissions - one agent runs as thousands of isolated versions. ### What surfaces can agents run on? Agent Chat SDK (ChatGPT/Claude-grade, voice, multimedia, canvas) or headless; in-app, web, mobile, Slack, Microsoft Teams; and proactively off your events. ### How is this different from a frontend framework like CopilotKit? A frontend framework gives you components to assemble the UI yourself, in code, one per tool. Mindset gives the whole harness (assemble, deploy multi-tenant, scale generative UI, improve from usage) and works with any framework. ### Do we have to use your models? No - model-agnostic; choose from available LLMs or bring your own. ### How do we know an agent is working, and how do we improve it? Observability dashboards capture every conversation; test with Preview and Test Cases; fix in config with no redeploy; export to your data lake or use built-in reports. ### What support do we get setting it up? Forward-deployed engineers build the first agents alongside your team to an agreed deadline and hand over the patterns. ## Take your agents from a few to a whole fleet - built, run and improved from one place. ## Finally, your team's skills, agents and MCPs are off their laptops. *For Heads of AI at mid-market B2B SaaS companies running internal agents at scale* Adoption exploded and everyone loves Skills, agents and MCPs, but now critical flows live on unauditable, unmonitorable or uneditable local devices. Mindset centralizes what your team built, and lets anyone, technical or not, build and improve it visually, on any model, independently audited, surfaced wherever your team works. Runs on the models you already use: Claude, OpenAI, Gemini. ## The rollout worked. Now business-critical AI runs the company, from individual laptops, owned by one person each. ### Trapped The skills and the connections live in one person's account. Nobody else can run them, change them or call them, so when they're away the business is too. ### Unguarded No evals or gates before a skill ships, and once it runs it reaches anything that account was handed a key to. No record of what it did or what it spent. ### Frozen Nothing is measured, so nobody knows where it's weak. Nothing is versioned, so a bad change can't be undone and a good one can't be repeated. ## Mindset is the agent harness & execution layer for your most important automations ## Everything your team built on laptops, rebuilt as automations the company runs, so it stops being shadow IT ### Turn skill.md docs on laptops into governed automations, with no engineering time required Mindset builds everything from a conversion or skill file. Connections, agents, if-this-then-that automation, acceptance criteria & more. ### Make your agents create reliable, improvable outputs every single week, swapping prompts and md docs for real automation. ### Re-prevision the governed automations into the tools your team works inside, so they don't need to learn a new app. ## Move automations onto Mindset, so they become governed, auditable & transparent ### See everything your team has built in one place, so nobody has to work out what is running All critical flows owned, versioned, visible in one place. People find what already exists and reuse it, instead of building the same thing for a third time. ### Decide exactly what your agents are allowed to touch, so you can minimise the blast radius. ### Prove what every agent did to any auditor in a click, so an audit is an export and not a project. ## Any model, any stack and self-hostable so you can stay optimized for the best of every technology ecosystem ### Any model, any ecosystem, self-hosted if you want it You stay in complete control; choose the best models, data never leaves your stack, you fit within the existing systems you have today. ### Continually optimize token costs, so you get the best results for the lowest price, instead of being blind to it. ### See exactly what’s breaking and improve every automation via a conversation with AI, so everything is easy to maintain. ## How it works ### Import & Build: Bring what you have in & give your team one place to build more. - Import the skills, agents, MCPs and automations off local machines into one org-scoped registry the company owns. - Build new ones or turn existing skills.md into deterministic flows if needed by describing what you want in plain language, including Functions that package complex business logic into a single tool your agents can call. - See the whole agent fleet in one catalogue: what's running, who built it, what it touches and what it costs, all versioned from the moment it lands. ### Govern & Deploy: Surface agents wherever they work & control who accesses what. - Runs on shared, monitored infrastructure with break alerts, independent of any person's machine or holiday. Your team decide what actions are recorded by speaking to AI. - Every agent gets its own identity and least-privilege access across tools and MCPs, instead of running on a human’s login. Agents and Functions can only touch the connections your org has sanctioned. - Route any agent, skill or flow to where people work already for the final output or human-in-the-loop approval. ### Improve: Turn ad-hoc flows into ones that suggest self-improvements you approve. - Set behaviour acceptance criteria per agent or flow. Mindset runs them before you publish and re-runs them on a schedule, so a flow that goes red tells you before a customer does. - A closed loop watches real usage and recommends changes (a tool called 27% of the time that should be nearer 50%), and lets you A/B test and simulate a change before it ships. - Token cost per skill shows what earns its keep and where to cut, and an independent audit trail on every run lets you scale to reputation-bearing work. ## Common questions ### Do we have to rebuild our skills? No. Mindset wraps the skills, agents and MCPs you already built and runs them as they are. ### Isn’t this just moving scripts onto a server? No, that's the tactical version. Mindset adds what a laptop can't: an independent audit trail, agent identity and least-privilege access, cost per skill, evals and quality gates. That's the governance that lets you scale to customer-facing work. ### How is this different from an AI governance dashboard like Arthur or Credo? Those watch and report on agents you still run elsewhere. Mindset runs them: registry, execution, audit and improvement in one platform, not observability bolted on top. ### How is a Function different from a Zapier or n8n workflow? Those follow fixed, deterministic steps, so they can’t make a judgment call ("does this payment match this invoice?") without you writing endless rules. A Mindset Function can include a real judgment step alongside the deterministic ones, is built by the person who owns the process in plain language, and runs on your own governed infrastructure. ### Does it lock us to one model or ecosystem? The opposite. It’s model- and ecosystem-agnostic, built on open protocols and exported under a perpetual source licence. Switch models whenever a better or cheaper option appears, and the IP stays yours. ### Do we need a platform team? No. Non-engineers build and run governed flows themselves; technical teams extend in code. Roughly a quarter of an engineer, and it keeps your senior engineers on IP, not maintenance. ### What about PII and data residency? You control what data reaches a model, with PII controls and documented residency: self-host or run in the US, UK or EU. ### Can we control what agents touch? Yes: agent identity instead of shared credentials, least-privilege access per agent, a full audit trail, and a kill switch. ### Where do approvals happen? In Slack or Notion, the tools your team already uses, not a new platform. ### How do we know a flow is good enough for customer-facing work? Evals and quality gates run before it ships, and an independent audit trail records what it did. ## Take control of the AI your team already built. ## Tell us what you're trying to achieve. *START HERE* A 30-minute introductory call. Share your goals, we show you the path — then we book the deep dive. - **Share your goals**: Tell us the KPIs you need to move, the business outcomes you care about, and where AI fits your roadmap. - **See how we can help**: We'll map your objectives to the platform — what Mindset handles, what stays yours, and what the timeline looks like. - **Get honest answers**: Not every problem needs our platform. We'll tell you where we add value and where you might not need us. - **Define next steps together**: If there's a fit, we'll book a technical deep dive with a solutions engineer scoped to your architecture. - **No commitment, no pressure**: This is the start of the conversation, not a sales pitch. Come with questions, leave with clarity. --- # Showcase The showcase has five areas: - **[Live Agents](https://www.mindset.ai/showcase/agents)**: In-app agents rendering generative UI in real time. See what your agents could look like in production. - **[Generative UI widget gallery](https://www.mindset.ai/showcase/widgets)**: Browse declarative UI widgets agents render at runtime — HTML components, charts, forms, tables, actions. Not just text. - **[Generative UI widget builder](https://www.mindset.ai/showcase/widget-builder)**: Build the generative UI components your agents render to users. Design visually in the studio, export as code, or build in your IDE. - **[Agent Management Studio](https://www.mindset.ai/showcase/agent-management-studio)**: The visual layer on top of the framework. Create agents via natural language prompts or pre-provided options. Build gen UI components. Observe performance, evaluate outcomes, improve without redeploying. - **[Agent Builder SDK](https://www.mindset.ai/showcase/agent-builder-sdk)**: Let your customers build and manage agents inside your product — no need to build administration tooling yourself. Full white-label branding, full tenant isolation, built on open standards. ## See what you can build with the Mindset stack Every agent below was built using the Mindset frontend stack and framework. Start a conversation with any of them — this is what you can build with our toolset. ### Mindset Guide An out the box, customisable customer experience agent. It doesn't just answer, it does the work. It takes people anywhere in a product, performs the task on their behalf, and represents your brand. Try it live: https://www.mindset.ai/showcase/mindset-guide ### Content Health Agent An admin agent embedded in a dashboard that diagnoses underperforming courses, drills into lesson-level engagement data, and recommends specific content fixes. See MCP tools, skills, widgets, situational awareness, page actions, and programmatic triggers working together. Try it live: https://www.mindset.ai/content-health-demo.html ### Customer Health Agent A CS manager's agent embedded in a portfolio health dashboard. It identifies at-risk accounts on a bubble chart, diagnoses compounding risk signals, drafts outreach to CSMs, and posts internal notes. See the agent see the page, act on it, and start working without being asked. Try it live: https://www.mindset.ai/customer-health-demo-v2.html ### Course Support Agent A learning agent that routes users through course discovery, prerequisite checks, and enrollment. See multi-step workflows, context retention, and conditional routing in action. Try it live: https://www.mindset.ai/showcase/course-support Suggested prompts: What courses are available for beginners?; Help me find a course on leadership; What are the prerequisites for the advanced track? ### Content Agent A knowledge agent that retrieves and synthesizes information from a content library in real time. See retrieval patterns, citations, and dynamic content rendering in action. Try it live: https://www.mindset.ai/showcase/content-agent Suggested prompts: What content do you have available?; Help me find something on AI agents; Summarize the key topics you cover ## Generative UI widget gallery Front-end components for every use case. Browse the gallery, use any widget as a base, and edit it in your IDE using the GenUI widget editor and builder. Each widget connects MCP tool outputs to visual MDX components. - **Chat Widget**: Full conversational interface with citations, media, and interactive elements. - **Card Widget**: Compact, embeddable cards that surface agent responses inline with your content. - **Mini Widget**: Minimal footprint widget for quick interactions without leaving the page. - **Banner Widget**: Full-width banners with contextual AI responses at the top of any page. - **FAB Widget**: Floating action button that expands into a full agent experience on tap. - **Inline Widget**: Inline agent responses that blend seamlessly into your existing page content. ### Widget Categories - Data Display - Interactive - Navigation - Media - Forms - Layout - Charts - Cards - Lists - Tables ### Generative UI widget builder Connect MCP tool outputs to visual MDX widgets. Create rich, interactive interfaces without code — or take the output into your IDE and make it yours. ## Let your customers build agents inside your product The Agent Builder SDK lets your customers build and manage agents directly inside your product — so you don't need to build endless administration and setup tooling yourself. Full white-label branding. Full tenant isolation. Built on open standards. - **Agent Builder**: Your customers create agents visually — define behavior, connect knowledge bases, configure tool access. - **White-Label**: Fully branded — your colors, your logo, your domain. The SDK feels native to your product. - **Multi-Tenant Isolation**: Each tenant's agents, data, and configuration are fully isolated. Security enforced at the framework level. - **Workflow Editor**: Your customers design structured workflows — onboarding, compliance, coaching — directly inside your product. - **Analytics Dashboard**: Built-in observability — question themes, sentiment, usage patterns. Your customers see how their agents perform. - **API-First**: Everything available via API. Automate agent management across your customer base programmatically. ## The management layer you don't have to build Every agent team needs a way to manage behavior, connect data, design workflows, and monitor performance. Building that management tooling yourself costs months. The studio gives your team a visual control center — update system prompts, adjust guardrails, iterate on agent behavior. Simpler, faster maintenance than managing it all in code. - **AI Agent Builder**: Define agent behavior — system prompts, guardrails, knowledge bases, and tool access. Build visually or control programmatically via API. - **Knowledge Management**: Manage the content sources your agents retrieve from. Upload, organize, and control access per tenant. - **Journey & Workflow Designer**: Design multi-step workflows visually. Define what happens at each stage — what the agent asks, what tools it calls, what triggers the next step. - **Widget Studio**: Connect MCP tool outputs to visual MDX widgets. Create rich, interactive interfaces without code. - **MCP Tool Integration**: Add, test, and manage MCP servers. Define what tools your agents can call and monitor performance. - **Analytics & Observability**: Monitor agent performance, conversation themes, sentiment, and usage patterns. Spot problems early, iterate, prove impact. ## Mindset Guide Every other guide just answers. Mindset Guide does the work. Mindset Guide is an in-product agent that takes people anywhere in your product, performs the task for them, and represents your brand. It works by voice or text, inside the product they already have. ### Four things no tour, tooltip or chatbot can do - **It acts**: it takes the action, not just describes it — it reads the screen, completes the task on the person's behalf inside their own logged-in session, so their permissions always apply. - **It goes anywhere**: it navigates across your entire product mid-conversation, highlights the exact thing someone is looking for, and points a live cursor at it. - **It represents your brand**: give it a unique personality, name, voice and animated character, so it looks and feels like a real part of your product. - **It improves**: tell it the outcome you care about and it works towards it, with sentiment analysis, performance analytics, and suggested config changes sent straight to you. ### Where it works Activation (PLG onboarding), conversion (DTC storefronts), time-to-task (admin-heavy software), and completion (learning and courses). It installs as an add-on to your existing product: add the `` element, configure it in the visual builder, and integrate via the APIs. --- # Solution Pages ## AI Agents for Learning Platforms *AI Agent Platform For Learning Platforms* ### Turn your learning platform into a conversational experience Ship AI agents that handle compliance chasing, bulk assignments, and manager escalations for admins. Learners get instant answers and practice scenarios through conversation. **Problem:** Learners don't have time for courses. Admins are drowning in configuration. The learning world is going conversational and you're falling behind. You've invested in AI cloud, data infrastructure, and your own team. You need something that accelerates them, not replaces them. Unlock legacy training content (SCORM) and turn static courses into dynamic, searchable knowledge inside learning agents. Basic Q&A chatbots are table stakes — customers expect GPT-like conversational learning experiences in your platform. **Solution:** Mindset AI makes it easy to add conversational learning experiences without rebuilding your platform. Connect to your existing APIs — no migration required, no vendor lock-in. Describe what learners and admins need — course Q&A, practice scenarios, compliance tracking, bulk training assignments — and we generate conversational agents connected to your stack. Enabling product, design, and engineering to collaborate and ship learning AI this quarter, not next year. #### Compliance Chaser Agent Identify overdue learners and send bulk reminders Admins waste hours manually identifying who's overdue on compliance training and chasing them individually. The Compliance Chaser Agent: - Identifies all learners overdue on required training automatically - Sends personalized bulk reminders with one click - Shows how long each learner has been overdue and their completion status #### Practice Scenario Agent Learners rehearse real-world situations through conversation Learners need to practice difficult conversations before real situations but have no safe space to rehearse. The Practice Scenario Agent: - Role plays realistic workplace scenarios with the learner in real-time - Provides feedback on tone, word choice, and approach after each response - Lets learners retry the conversation with different strategies #### Bulk Assignment Agent Roll out mandatory training to entire teams in one conversation Admins need to assign new courses to hundreds of people but spend time manually selecting learners, setting deadlines, and sending notifications. The Bulk Assignment Agent: - Assigns courses to entire departments or cohorts with a single prompt - Sets deadlines and sends automated enrollment notifications - Shows confirmation of who was assigned and when they need to complete #### Coaching Agent Answer questions, provide Socratic coaching, and search across SCORM content Learners remember concepts from training but can't find the details across scattered SCORM courses, so they interrupt managers or don't apply what they learned. The Coaching Agent: - Searches and digests SCORM content from 5+ major providers to answer questions with specific explanations - Delivers quick assessments to check understanding and reinforce learning - Provides Socratic coaching by asking guiding questions that help learners think through problems themselves #### Manager Escalation Agent Escalate overdue training to managers when reminders fail Admins have chased overdue learners multiple times, but some still haven't completed mandatory training — now it's time to escalate to their managers. The Manager Escalation Agent: - Identifies learners significantly overdue and groups them by manager - Drafts escalation emails to managers with overdue report details - Sends notifications and tracks manager follow-up ## AI Agents for HR & People Ops *AI Agent Platform For HR & People Tech* ### Turn your HR platform into a conversational employee experience Ship recruiting agents, HR service agents, coaches, and onboarding assistants this quarter without depleting engineering resources. **Problem:** HR admins are drowning in manual work and your product feels behind. The HR world went conversational and you're not keeping up. You want to ship AI features now, but your engineering team is stretched across 10+ priorities building foundational infrastructure. Starting custom AI builds from scratch means shipping next year while competitors launch in weeks. HR admins expect ChatGPT-quality AI that answers policy questions, checks leave coverage, and automates approvals. Your click-through menus and manual workflows feel like yesterday's software. **Solution:** Mindset AI makes it easy to ship conversational HR agents without rebuilding your platform. We plug into your existing infrastructure via APIs. Everything you build stays yours and exports as code. Tell us what HR admins and employees need — policy answers, leave approvals, employee look-ups, onboarding workflows — and we generate agents that connect directly to your APIs. Your product, design, and engineering teams work together to ship AI features fast while keeping full control. #### HR Admin Agent Handle employee look-ups, absence checks, and policy questions in seconds HR admins answer the same questions dozens of times a day — remaining holiday allowance, who's out this week, policy eligibility. The HR Admin Agent: - Surfaces employee details and holiday allowance instantly - Shows team availability and flags coverage gaps before approving requests - Answers policy eligibility questions based on each employee's tenure and contract #### Compliance Management Agent Find compliance gaps, chase outstanding training, and stay audit ready Tracking who's completed mandatory training or missing certifications usually means spreadsheets and manual follow-ups. The Compliance Agent: - Identifies gaps across your workforce - Lets you assign employees from a conversation - Gives you a complete audit view before regulators arrive #### Absence & Coverage Agent Check team coverage, approve leave requests, and avoid scheduling conflicts Managers ask if they can approve a leave request — admins need to check who else is out before saying yes. The Absence & Coverage Agent: - Shows real-time team availability for any date range - Flags coverage risks and warns when approvals drop below minimum staffing - Helps admins make confident approval decisions in seconds #### Employee Lifecycle Agent Automate onboarding, exits, probation reviews, and leave returns HR admins manually coordinate every lifecycle transition — new hires, exits, probation reviews, and returns from leave. The Employee Lifecycle Agent: - Kicks off onboarding tasks for IT, facilities, and managers automatically - Processes exits by revoking access and scheduling handovers - Confirms probations and triggers review workflows based on start dates - Handles leave returns by notifying managers and reinstating access ## AI Agents for Talent Acquisition *AI Agent Platform For Talent Acquisition & Recruitment* ### Turn your talent acquisition platform into a conversational recruiting engine Ship interview coaches, scheduling agents, and outreach assistants this quarter without depleting engineering resources. **Problem:** Recruiters don't have time for manual coordination. Admins are drowning in configuration. The recruiting world is going conversational and you're falling behind. Your AI team is building core platform capabilities while the recruiting product roadmap stacks up. You can't pull engineers off infrastructure to build custom recruiting agents. Recruiters drown in manual admin — resume parsing, interview scheduling, candidate follow-ups eat 20+ hours per week. Customers expect ChatGPT-quality AI that understands their job descriptions, candidate pools, and hiring workflows. **Solution:** Mindset AI makes it easy to add conversational experiences to your recruiting platform without rebuilding it. We integrate with your existing stack via APIs. No migration, no vendor lock-in. Everything you build remains your IP and exports as code. Describe what users need in conversation — candidate outreach, interview scheduling, feedback collection — all connected to your APIs. #### Scheduling Agent Coordinate panel interviews across multiple calendars in seconds Recruiters coordinating panel interviews manually check 3-5 calendars, send availability requests, and play email ping-pong for days. The Scheduling Agent: - Finds available time slots across all interviewer calendars instantly - Shows candidate availability and suggests optimal times - Books the interview and sends calendar invites with one click #### Outreach Drafter Agent Draft personalized candidate outreach in seconds, not hours Recruiters sourcing passive candidates spend hours crafting compelling, personalized messages for each prospect. The Outreach Drafter Agent: - Drafts personalized outreach tailored to candidate's background and specific role - Pulls in relevant details from LinkedIn profiles and job descriptions automatically - Generates messages that sound human, not templated #### Feedback Chaser Agent Stop chasing interviewers for feedback — automate follow-ups instead Recruiters waste hours chasing interviewers who haven't submitted feedback, blocking hiring decisions while candidates wait. The Feedback Chaser Agent: - Identifies all overdue feedback from the past week automatically - Sends personalized reminder messages to interviewers - Shows who's blocking each decision and how long candidates have been waiting #### Rejection Sender Agent Close out candidates respectfully without manual follow-up Recruiters need to send rejection emails to dozens of candidates stuck in screening, but writing personalized, respectful messages takes time they don't have. The Rejection Sender Agent: - Identifies all candidates who've been in screening longer than 3 weeks - Generates respectful, personalized rejection emails at scale - Sends bulk rejections while adhering to employer brand guidelines ## Conversational AI Agents for SaaS *Conversational AI Agent For SaaS Products* ### Make your SaaS conversational with AI-powered agents Your users navigate menus, fill forms, and click through outdated UI flows. Give them the faster way they expect: ask what they need, and AI agents execute. Mindset AI is the conversational experience layer that makes it happen. **Problem:** Your product feels like legacy software because it's not conversational. 1.5 billion users interact with AI apps daily. When they come to your product and navigate through menus and forms, the experience feels outdated. Leadership needs it done yesterday — the board, C-suite, and sales expect agentic AI to increase engagement and reduce support costs this quarter. But you need more than a chatbot — widgets, actions, multi-channel, memory, and multi-tenant deployment. **Solution:** Build AI agents that control your product's features. Connect them to your data and systems, so they can act, not just chat. Configure agents visually, deploy a production-ready chat UI, wrap your APIs as MCP servers, and build interactive widgets that surface during conversations — all without rebuilding your platform. #### AI Agent Builder Configure AI agents visually with guardrails and memory Subject matter experts configure agent behavior — guardrails, personality, and memory settings — on a visual canvas. Engineers control which tools agents can access by wrapping APIs as MCP servers and managing deployment across tenants. Prototype and ship fast with both teams contributing their expertise. #### Agent Chat SDK Deploy a production-ready chat UI for agents, tools, and widgets Launch agents via a Perplexity-like chat experience directly in your product. Add with built-in search, citations, suggestions, media, and widgets. Choose from sidebars, trays, or full chat layouts — all fully customizable. #### Connect Your APIs Enable agents to control your product's features Wrap your APIs as MCP servers so agents can execute actions — list_users, add_course, get_course_details. Build once, share across all agents. The standardized format means no rework as you add capabilities. #### Interactive Widgets & Widget Builder Build custom widgets that surface during conversations 'Book me a demo' — a calendar appears instantly. 'Update my subscription' — a plan selector renders in conversation. 'Show Q3 performance' — a chart displays. Build widgets with AI, code, or upload existing components. Connect to MCP servers for live data and actions. #### Agentic Workflow Builder Build structured journeys for AI agents to guide users through Some user journeys need more than answers — guidance through a specific process. Onboarding, compliance walkthroughs, coaching sessions, and approval routing. Design exactly how the agent moves users through each step. #### Agent Memory Agents capture what matters about each user to personalize future interactions Agents automatically extract facts from conversations — job role, communication preferences, current projects, goals. These facts build a profile that makes every interaction more relevant without manual configuration. #### Agent Sessions & Content Ingestion API Control which AI agents and actions users access via API Deploy agents across all tenants with full control over who can access what content, tools, and workflows. Manage permissions at scale through APIs, keeping control in your developers' hands. ## Build & Monetize AI Agents *Build, Ship & Monetize AI Agents This Quarter* ### Build and monetize AI agents that take action — this quarter, not next year Build agents visually that take action and create real value, then deploy across thousands of tenants with agent memory, enterprise guardrails, and multi-tenant APIs built in. **Problem:** Users pay for agents that execute actions. Your team can build the AI logic but get blocked on the UI infrastructure layer. Product teams are blocked — subject matter experts know what customers need, but wait months for engineering resources to build and ship agentic features. UI work doesn't differentiate — engineers should focus on AI that sets your product apart, not building UI widgets, chat interfaces, and compliance infrastructure. Sales cycles extend indefinitely as enterprise buyers demand security, compliance, and visibility before signing. **Solution:** Build action-taking agents visually, deploy across all tenants programmatically, and ship with enterprise security built in. Configure agents on a visual canvas with guardrails, personality, and memory. Wrap your APIs as MCP servers so agents can execute real actions. Deploy across thousands of tenants via API with per-customer customization. Enterprise security, compliance, and audit logging are built in from day one. #### AI Agent Builder Configure agents visually with guardrails and memory Build agents visually — guardrails, personality, memory settings — on a visual canvas. Engineers control which tools agents can access by wrapping APIs as MCP servers and managing deployment across tenants. Both teams contribute their expertise to ship fast. #### Agent Chat SDK & Widget Studio Surface interactive components directly in chat to differentiate your offering Provide visual and interactive experiences with UI widgets that render in chats for users — a unique experience people will pay for. - 'Book me a demo' — a calendar appears for the user to pick a suitable slot - 'Update my subscription' — available plans render in conversation - 'Show Q3 performance' — data displays in a chart #### Connect Your APIs Enable agents to control your product's features Wrap your APIs as MCP servers so agents can execute actions — list_users, add_course, get_course_details. Build once, share across all agents. The standardized format means no rework as you add capabilities. #### Agent Sessions & Content Ingestion API Scale one agent across thousands of tenants with customization per customer Deploy agents via APIs across all tenants with full control over access permissions. Let customers customize agents or upload their own content. One agent becomes thousands of customized versions — every tenant gets a personalized experience, all managed through APIs. #### Agentic Workflows Build agent-led journeys and control access for monetization Design structured processes for onboarding, compliance, coaching, approvals and more. Agents guide users through each step of the journey. Publish workflows and control access via API for monetization. #### Enterprise Security Deploy with enterprise security built in Enterprise customers want AI agents but fear security risks. Built-in guardrails, policy controls, audit logs, reporting, ISO certification, and compliance with global AI frameworks. Sell into regulated industries confidently. ## Snowflake Agent Acceleration *Snowflake Agent Acceleration* ### Snowflake-powered? Ship customer-facing AI agents in weeks, not months You've got rationalised data, strong models, and serious infrastructure. Mindset AI is the frontend layer turning that into something your customers can actually use. Build APIs on your Snowflake data, expose them to agents through MCP, and ship AI experiences where your customers talk to your product and take real actions. **Problem:** You've invested heavily in Snowflake, but you're not yet shipping customer-facing AI on it. Your engineers would need to build chat interfaces, widget systems, and front-end deployment infrastructure from scratch. None of it is what makes your product unique. Every new capability means weeks of frontend development — even when the business logic already exists. Your data and ML teams build powerful differentiation, but getting it into customer hands requires coordinating across product, design, and engineering, slowing time to value. **Solution:** Mindset AI connects to your services and business logic through MCP and gives your customers a conversational way to interact with all of it. Build on Snowflake. Then ship conversational experiences in weeks, not quarters. Mindset AI handles the frontend — you focus on building what makes your product unique. #### Agent Chat SDK Add AI-native chat experiences that your customers can access Deploy a Perplexity-like agent chat directly in your product. Add with built-in search, citations, suggestions, media, and widgets. Agents & widgets you build surface inside the chat for users to access. #### Interactive Widgets Widgets complete customer tasks in one interaction, not ten clicks Every new customer-facing capability needs a frontend build. Weeks of work, every time. With Mindset AI: - Skip the frontend work that doesn't differentiate you — build widgets connected to your APIs via MCP - Agents render interactive components in conversation (book a meeting, update subscription, assign users) #### Connect Your APIs via MCP Connect your platform and services. Let agents do the rest. You've built services, models, and APIs on Snowflake. But each new capability means weeks of frontend development before customers see it. With Mindset AI: - Expose your APIs (recommendation engines, pricing logic, scheduling services) as tool-based MCPs - Connect existing text-to-SQL tools as MCP servers (your own, dbt, or Snowflake Cortex) - Agents reason and decide which APIs to call for users - Widgets pull live data and render in context — calendars, forms, data comparisons - We never access your databases directly — everything flows through MCP services you control #### Turn ML Models into Customer-Facing Experiences Surface predictions and recommendations through conversation Your teams have built ML models that generate predictions and recommendations. Turning them into something customers can use is a separate project entirely. With Mindset AI: - Agents tap into your ML models through MCP services — so questions like 'Who\'s at risk of churn?' just work - Predictions surface as interactive widgets — visual indicators, trend charts, and key drivers behind them #### Enterprise AI Features Deploy production & enterprise-ready agents with compliance built in Production-grade AI. Enterprise-grade governance. With Mindset AI: - Multi-tenant provisioning with per-customer data isolation - Agents connect through MCP services you expose — no direct database access - Compliance guardrails and built-in audit logging - Memory engine personalizes each experience within your data boundaries ## Optimize ### Have total confidence your agents work, improve them through conversations & optimize LLM costs continually Mindset AI tracks real user usage, assesses against your agent acceptance criteria, identifies failures, lets you interrogate the data & improve agents via conversations #### Cost and ROI: Continually reduce token costs, without sacrificing automation outputs Every model call is metered and attributed at the point it happens, so the bill is explicable before anyone asks. - **Cost breakdown** Spend rolls up by agent, user, model, provider and org, so your bill finally names names. - **Cost analyst** Ask in plain language where the money went, backed by spend over time and by agent. - **Model cost comparison** Token-level capture prices what that agent would have cost on a cheaper model, before you switch. #### Behavior testing and evals: Catch a bad answer before your users do with simulations & acceptance criteria. You state the requirement and scenario in plain language and the platform turns it into criteria it can score. - **Acceptance criteria** Say what good looks like in plain language and the builder turns it into agent acceptance criteria checks. - **Real agent testing:** Tests run the real agent your customers use against your criteria, flagging problems & suggesting fixes. - **Side-by-side model tests:** Run the same tests against two models side by side, and see which answers better so you pay less for the same outcomes #### Observability: See exactly what's breaking instead of being frustrated & lost One board ranks the whole estate of your tools, agents, LLMs, agent delegation and more, so you can find the problem instantly. - **Health board** Agents, functions, MCP servers, agent reasoning, every tool and more, ordered by what is failing now - **Tool health summary** Click any tool to see how often it ran, how often it failed, and how slow it was - **OpenTelemetry export** The same spans leave over OTLP, into whatever your team already audits and retains. #### Traces and replay: Replay any run and see exactly why it did that. A run is a pure function of three recorded inputs, which is what makes replay exact rather than approximate. - **Replay on demand** Run the same conversation again, from the same starting point, and watch it fail identically. - **The whole decision** Every step the agent took, every tool it called, and everything those tools returned. - **Prove the fix** Change the prompt, replay the same case, and see whether the behaviour actually moved. #### Estate world graph: See how your whole estate fits together. One live graph of the whole org, with an agent that reads it for you, watching live as everything is executing. - **Whole-org graph** Agents, tools, functions, content banks and more in one view, so you can visualize what's connected. - **Guide agent** The agent helps you delve deeper into every use-case and automation - **Shadow AI discovery** Flows your people built themselves arrive in the same graph, with what data they touch and cost. Know your agents work before a person tells you they don't. ## AI Control Plane ### Run and govern the agents, skills and MCPs your team already built Move them off personal laptops and run them where you can see them, improve them, control what they cost and prove what they did. Any model. Your servers or ours. #### The catalogue: See everything your team has built, in one place Every skill, agent, function, connection and MCP server your team builds lands in one catalogue the company owns. - **One catalogue** Every agent, skill, function, connection and MCP your team builds lands here. - **Owner and version** Every one shows who built it, what version it is, and what it touches. - **A live graph** See what is connected to what. You do not have to work it out. #### Permissions: Give someone an agent, and you have given them their access The agent is the permission. Every agent carries the connections, tools and data it can reach, so granting someone that agent grants exactly that and nothing else. - **Grant the agent** Give someone an agent and they get everything that agent can reach. Nothing more. - **On the agent** Connections, tools and data are attached to the agent, not to the person. - **Works in Claude** Anything that supports MCP shows only the agents that person may use. #### Approved connections and MCPs: Decide exactly what your agents are allowed to touch Register the APIs and MCP servers your agents may use. Each one gets its own login with only the access it needs. - **One way out** Every call to an outside system goes through a connection you approved. No exceptions. - **Approved only** Agents can only use the connections and MCP servers you have registered. - **Company logins** Access belongs to the company, not a person. Nothing breaks when someone leaves. #### Personal data: Keep personal data away from every model Personal data is swapped for realistic fake data before a prompt reaches a model, then swapped back in the reply. - **Fake data in** Personal data is replaced with realistic fake data before the model sees it. - **Real data back** The real values are put back into the reply the user reads. - **Your rules** You decide what counts as personal data. We do not decide it for you. #### Audit trail: Prove what every agent did, to any auditor Every agent action is recorded as it happens. Export a compliance-ready report for EU AI Act, ISO 27001, SOC 2 or GDPR in one click. - **Everything is logged** Every action an agent takes is recorded. Agents cannot skip the record. - **Nobody edits it** Nothing is ever deleted or rewritten. Not by your team, not by ours. - **Export a report** Get a compliance-ready report in one click. Weeks of collating, gone. #### Versions and rollback: Change any agent, and roll it back in one click Every change to an agent is saved as a version. A change only reaches real users when you publish it. - **Every change saved** Prompts, tools, connections and settings are all versioned. You can see what changed. - **Roll back instantly** Made a change that broke something? Put the old version back in one click. - **Drafts stay drafts** A change only reaches real users when you choose to publish it. #### Where it runs: Run the whole platform on infrastructure you own Run it on your own servers from day one, or pick the US, UK or EU. Runs pick up again after a failure either way. - **Your own servers** Run the whole platform in your own cloud. Or use ours. - **Pick your region** Hosted in the US, UK or EU. Your data stays where you need it. - **Any model** Switch models per agent. Put the cheap model on the cheap work. #### Export to your own tools: Send your agent data into the tools you already run Agent activity streams out as OpenTelemetry, so it lands in Datadog, Splunk or whatever your team already uses. - **No new tool** Your team keeps using Datadog or Splunk. Nothing new to learn. - **Your tags included** Your own tags and user IDs come through, so you can filter by customer. - **You choose what** We only send what you switch on. Nothing leaves without you saying so. Take control of the AI your team already built ## Channels ### Put your agents where people already work Put your automations into the places your team work, respecting your existing permissions #### In your team's AI tools: Call every automation from ChatGPT, Claude, Gemini & more Your governed agents show up as callable, audited tools inside whatever AI your team already uses. - **Connect once** Every automation you've published becomes callable inside any MCP-supported system your team already uses. - **Scoped access** Nobody reaches an agent their permissions block, and every call is recorded for compliance. - **Any AI tool** ChatGPT, Claude, Gemini or your own. Adding a new one needs no extra work. #### Slack & Teams: Run agents in a thread, publish results to a channel Chat with a governed agent and get its results right inside the tools your team already works in, Slack and Teams, with no new dashboard to log into. - **In-thread answers** Ask in a channel and the result comes back in that thread, not behind a link. - **One configuration** Slack & Teams are a surface, not an ungoverned side channel with its own permissions to maintain. - **Full audit trail** Every question asked in a channel is recorded exactly as one asked in the console. #### Web page or product: Add a ChatGPT-grade chat to any page without building one A ChatGPT/Claude-grade chat (voice, multimedia, canvas) that you host as a ready-made page or drop onto any site with a snippet or an iframe. A familiar, high-quality interface for your team or users from day one. - **Snippet install** One script tag and a element puts the full chat on any page. - **Visible or headless** Commands in, typed events out, so you can render the agent in your own UI. - **Feature-rich** Voice, multimedia, citations, canvas, Gen UI widgets and more out the box Meet people where they already are ## SDKs & APIs ### Developer tools for building agents throughout your product Mindset is the harness for agents your customers use - build from composable blocks, improving from real usage, fixes you approve in conversation, not tickets. #### Embed agent building: Let your customers build their own agent automations Every screen we provide for your team is a component you can put in your own product, or replace entirely with your own. - **Embeddable builder** The UI your team uses to build agents, functions, widgets, and connections is components you can drop straight into your product for your users - **Your brand:** With design tokens, make every component look exactly like your own product - **Headless:** Build your own UI on top of any component, so you can control the whole experience yourself and run Mindset underneath it #### Agent chat: Drop in our chat, or build your own One package covers both, and moving from one to the other later does not mean starting again. - **Ready-made chat** Integrate a ChatGPT & Claude-grade chat experience. Citations, canvas, Gen UI widgets, media, thought process, tool UI, thread management & more. - **Headless:** Integrate agents into your UI, chat, or build your own chat interface on top. - **A published contract:** You get a contract, so anything that would break you arrives as a new version you choose to take, on your release schedule rather than ours. #### Widget Studio: Build generative UI widgets agents render in conversations The conversation stops being a wall of text and starts being something the user can actually use. - **Widgets inside the conversation** Cards, forms, comparisons, and charts appear in the chat at the exact point the agent decides they would help, rather than the agent describing something the user then has to go and find. - **Built from your components** Widgets are authored on your own design system, so they look like your product - **Build in Mindset or via MCP:** Use our MCP so you can author from your own IDE, or inside the Agent Studio #### Situational awareness & page tools: Your agents can see, control & navigate across your UI Any action in your front end can become something the agent sees, navigates users to or chooses to do on the user's behalf. - **Live page context** The agent runs inside your page, so the record they have open, the fields they have filled and the state around them are already available to it. - **Page tools** Register an action once and it becomes something the agent can use: open a modal, prefill a field, highlight a button, move someone to the right screen. - **Nothing happens that you did not allow** The agent can only use actions and values you registered. #### Agent sessions: Build one agent and offer unique experiences across thousands of users, not an agent per customer You maintain a single agent. Every person who opens it gets a private version, built for them the moment they start talking. - **One agent, every customer:** You build and improve one agent. There is no forked copy per customer, no separate deployment per account, and no rebuild when you sign the next one. - **Each session belongs to one person** The moment someone starts a conversation, their agent is assembled for them: their content, tools & more. - **You decide who gets what** Your system already knows what each user is allowed to see, and tells us at the start of the session. #### Build via MCP: Build everything via your IDE Via MCP, add the Mindset Admin agent builder into your Claude Code, Codex, Copilot or Cursor with one command, and your coding agent can build, change and test agents the same way your team does in the studio. - **Your IDE:** Agents get built where you already work, by the agent you already use. - **Drive it from your own code** The same operations are a documented API, so agent management can live in your provisioning scripts, your CI, or your own admin panel. - **Secure** A change made from an IDE goes through the same permissions as a change made by a person, and lands in the same record. Put a governed agent inside your product. ## Agent Studio ### Build powerful automations in Mindsets visual studio Turn an idea, spec or existing skill.md into agents with tools, if-this-then-that automations and sub agents by speaking to AI #### Skill.md ingestion: Turn skills on laptops into governed automations - by speaking to our AI Tell Mindset AI what you want to build or give it your skill.md files & MCPs your team already built and turn them into governed, reliable automations. - **Skills and MCPs** Hand over what your team wrote and get back an agent with its functions and connections. - **Approved plans** You get a build plan first: agents, functions, connections. Nothing is created until you say yes. - **Reuse first** It checks what you already have, so you don't end up with four copies of one connection. #### Agent Connections & Integration hub: Connect & govern the systems your agent needs to reach Set up new connections in one place. Every system your agent reaches goes through it, where you can see what it can do, approve it and take it away. - **Connection types** Connect to 1000s of agent tools or add your own MCP servers, HTTP APIs, Postgres, Google Sheets, knowledge sources. - **Live-tested setup** Our connections builder agent sets up your MCP, APIs & more for you, using documentation, URLs or whatever you provide it with - **Approvals and secrets** A person approves anything with side effects, and credentials never reach the model. #### Functions: Give an agent an if-this-then-that workflow by speaking with AI Calculations, data transformations, lookups, calls into your systems that stay the same every run. The agent decides when to use it or you force it to use it under specific circumstances. - **Built by AI** Describe the automation in natural language, plus the connections it needs. AI builds it for you. - **Deterministic steps** Calculations, lookups and transformations run identically every time, and never call the model. - **Model steps** Add a model call only where you want judgement, like drafting a reply. #### Multi-agent orchestration: Split the work across specialist agents A conductor agent delegates to a team of specialists, so complex work can be handled by focused agents that each do one thing well, instead of one over-loaded prompt. - **Any agent can be a resource** Attach an agent to another agent the same way you'd attach a function or a connection. - **Focused agents** Each specialist can carry simple & short instructions instead of one giant prompt and tool set, increasing performance & reducing LLM costs. - **Permissions** A specialist runs with the access of whoever asked so responses are always permission based, and every hand-off is recorded. #### Agent Scripts: Give agents a process that makes it consistent & dependable Without a script, an agent is a prompt and tools, and you're trusting the model to follow your instructions. It will 'probably' follow them; that's it. A script makes it dependable & cost-effective. - **Phases** Break the job into ordered phases, each with a goal and only the tools, agents & functions it needs to complete the job. - **Gates** A condition it must pass to move on. Until then it keeps working the user toward it. - **Provable behaviour** Write down how it must behave, test it, and guarantee dependability. #### Widget Studio: Design rich UI your agent shows in chat Someone asks to book a demo and a calendar appears. Someone asks to change plan and a plan selector renders in the chat. - **Interactive** Someone picks a date or a plan in the chat, and the agent acts on it. - **Live data** Widgets read from your connected systems, so people see what is true right now. - **Visual guides** Or use a card or chart simply to make a long answer readable at a glance. Build your first governed agent today. --- # Platform Overview ## The solution ### Mindset is the platform to build, run and govern the AI your team already built. #### Skill.md ingestion: Turn skills on laptops into governed automations - by speaking to our AI Tell Mindset AI what you want to build or give it your skill.md files & MCPs your team already built and turn them into governed, reliable automations. - **Skills and MCPs** Hand over what your team wrote and get back an agent with its functions and connections. - **Approved plans** You get a build plan first: agents, functions, connections. Nothing is created until you say yes. - **Reuse first** It checks what you already have, so you don't end up with four copies of one connection. #### Agent Connections & Integration hub: Connect & govern the systems your agent needs to reach Set up new connections in one place. Every system your agent reaches goes through it, where you can see what it can do, approve it and take it away. - **Connection types** Connect to 1000s of agent tools or add your own MCP servers, HTTP APIs, Postgres, Google Sheets, knowledge sources. - **Live-tested setup** Our connections builder agent sets up your MCP, APIs & more for you, using documentation, URLs or whatever you provide it with - **Approvals and secrets** A person approves anything with side effects, and credentials never reach the model. #### Functions: Give an agent an if-this-then-that workflow by speaking with AI Calculations, data transformations, lookups, calls into your systems that stay the same every run. The agent decides when to use it or you force it to use it under specific circumstances. - **Built by AI** Describe the automation in natural language, plus the connections it needs. AI builds it for you. - **Deterministic steps** Calculations, lookups and transformations run identically every time, and never call the model. - **Model steps** Add a model call only where you want judgement, like drafting a reply. #### Behavior testing and evals: Catch a bad answer before your users do with simulations & acceptance criteria. You state the requirement and scenario in plain language and the platform turns it into criteria it can score. - **Acceptance criteria** Say what good looks like in plain language and the builder turns it into agent acceptance criteria checks. - **Real agent testing:** Tests run the real agent your customers use against your criteria, flagging problems & suggesting fixes. - **Side-by-side model tests:** Run the same tests against two models side by side, and see which answers better so you pay less for the same outcomes #### Permissions: Give someone an agent, and you have given them their access The agent is the permission. Every agent carries the connections, tools and data it can reach, so granting someone that agent grants exactly that and nothing else. - **Grant the agent** Give someone an agent and they get everything that agent can reach. Nothing more. - **On the agent** Connections, tools and data are attached to the agent, not to the person. - **Works in Claude** Anything that supports MCP shows only the agents that person may use. #### Personal data: Keep personal data away from every model Personal data is swapped for realistic fake data before a prompt reaches a model, then swapped back in the reply. - **Fake data in** Personal data is replaced with realistic fake data before the model sees it. - **Real data back** The real values are put back into the reply the user reads. - **Your rules** You decide what counts as personal data. We do not decide it for you. #### Observability: See exactly what's breaking instead of being frustrated & lost One board ranks the whole estate of your tools, agents, LLMs, agent delegation and more, so you can find the problem instantly. - **Health board** Agents, functions, MCP servers, agent reasoning, every tool and more, ordered by what is failing now - **Tool health summary** Click any tool to see how often it ran, how often it failed, and how slow it was - **OpenTelemetry export** The same spans leave over OTLP, into whatever your team already audits and retains. #### Cost and ROI: Continually reduce token costs, without sacrificing automation outputs Every model call is metered and attributed at the point it happens, so the bill is explicable before anyone asks. - **Cost breakdown** Spend rolls up by agent, user, model, provider and org, so your bill finally names names. - **Cost analyst** Ask in plain language where the money went, backed by spend over time and by agent. - **Model cost comparison** Token-level capture prices what that agent would have cost on a cheaper model, before you switch. See the whole platform, running on what your team already built. ## Technical Overview *DEEP DIVE* ### Technical Overview of the Mindset AI Platform A comprehensive technical walkthrough of the Mindset AI platform — from the three-layer architecture and production components to SDK integration, enterprise MCP connectivity, and the 10 key differentiators that set Mindset apart. #### Generative UI is the new frontend For decades custom UI was the competitive advantage. That era is ending. Customers expect interfaces that generate components on demand, render voice, stream agent reasoning, and support human-in-the-loop workflows. The interface paradigm has shifted from static to generative. #### The Three-Layer Architecture The modern agent architecture has three layers. Layer 3: your data and infrastructure. Layer 2: your APIs, business logic, and domain expertise — your differentiation. Layer 1: the frontend stack for generative UI agents. Building Layer 1 yourself takes months. Using Mindset means days to production. #### What the stack gives you Mindset is the framework layer for generative UI agents. You bring the business logic and APIs. Mindset gives you: generative UI rendering, tool calling, widget systems, voice, streaming agent processes, human-in-the-loop controls, session management, multi-surface deployment, MCP integration, compliance (GDPR, CCPA, EU AI Act), observability and testing. You bring: APIs, data, business logic, LLM choice, RAG, domain expertise, user authentication. #### SDK & API Integration Three paths to production — Chat SDK: add in 10 lines, fork it, or bring your own client. REST API: full programmatic control over agents, sessions, memory, and observability. Headless Agent API (gRPC): streaming agent reasoning, state management, complete control. All paths give you the same generative UI, voice, widgets, and observability. #### MCP: the tool layer Use MCP — an open standard — to define the APIs your agents call. Add APIs, manage schemas and authentication, test agent tool calls, debug interactions. Data stays in your systems — agents call your APIs in real time. Open standard, not proprietary. #### Why developers choose Mindset's stack Why developers choose Mindset's stack — (1) Generative UI: agents render HTML components contextually — not just chat (2) Voice: agents speak and listen across channels (3) Streamed reasoning: users see the agent thinking in real time (4) Human-in-the-loop: agents pause for approval on high-stakes actions (5) Shared state: bidirectional communication between agent and UI (6) Component export: React and Web Components — use anywhere, no lock-in (7) MCP-first: all integrations use an open standard (8) Multi-tenant from day one: session-level isolation, compliance, memory (9) Observability: monitor agent behavior, tool calls, and performance at scale (10) Interoperable: built on LangChain, runs client-side, open standards throughout ## TL;DR *TL;DR* ### In 30 seconds: build agents faster, with confidence The frontend stack for in-app agents with generative UI. Go from your APIs to production in days — without sacrificing control. #### What Is Mindset AI The frontend stack for building in-app agents. - Agents that render generative UI components on demand — forms, charts, approvals, custom HTML - Interactive widgets triggered by agent tool calls via MCP - Deploy across web, mobile, Slack, Teams, and more #### How It Works Start building in days, not months. - Chat SDK — 10 lines of code, fork it, or bring your own client - Connect your APIs via MCP — an open standard - Your data stays in your systems; agents query it in real time #### Why Mindset Speed. Confidence. Control. - Production-grade from day one — multi-tenant deployment with session-level isolation - Compliance built into the framework: GDPR, CCPA, EU AI Act - Built on open standards — fork the SDK, bring your own LLM, export everything as code - Observability and automated testing included #### Build your first agent this quarter Build your first agent this quarter. - Visual agent builder — configure system prompts and tool access, no code needed - Agent Builder SDK — your customers create agents inside your product - SDK + quickstart gets you to prototype in hours, production in days --- # Browser-Based Agent Framework ## Your customer-facing agents are slow. You are paying a stateless tax. *For engineering leaders* Mindset AI is the browser-based agent framework that fixes it. Speed at depth, page-aware, and behaviour you can change without a deploy. Watch our Chief Product Officer Jack and our AI R&D lead Wic walk through the whole shift in our new webinar. Get it in your inbox in minutes. ### Problem You've added more tools to your agent to make it smarter. By turn nine the payload has grown from 2KB to 38KB, 82% of it redundant, and your users are waiting 6-12 seconds. And your agents cannot see the web page they're embedded in. ### An ultra-fast agent framework in the browser. Mindset gives you an ultra fast agent framework in a browser, with a visual builder on top for system prompts, adding MCPs, widgets, skills and more. #### Fast at turn nine. Same as turn one. The reasoning loop runs in the browser, so context doesn't rebuild on every turn. First-token latency holds at 1-2 seconds with 20+ tools active, however deep the conversation goes. #### Reads the page. Finishes the task. It pre-fills forms, applies filters, and completes the actions users were about to do manually. Five primitives give the agent developer-controlled access to page state and the DOM. #### Change behaviour without a deploy. Agent behaviour is versioned markdown, not code. Product teams ship changes in the studio in minutes, not release cycles. Engineering sets entitlements once, scoped per tenant. ### Client holds state. Server stays a proxy. Most agent frameworks today run the agent loop on the server. Every HTTP request lands fresh, with no idea who the user is, which thread they're in, which tools have already been called, or what the model said last turn. To work around that, the framework rebuilds the full thread, tool schemas, and memory blob on every request. The model gets it, the server forgets it, the next turn ships it again. Mindset AI moves the loop into the browser. The server only handles what genuinely needs the server. #### Step 01 — Add the embed. A single web component on your page. No backend wiring, no new endpoints, no deploy step on your side. #### Step 02 — Integrate the agent into the page. Five front-end primitives expose page state and DOM access, with the developer in control of what's exposed. The agent reads what's on the page and acts on it where you allow. #### Step 03 — Configure agent behaviour in markdown. Prompts, plan steps, policies, tools: all versioned markdown in the agent builder. Product teams iterate without a release cycle. Engineering sets entitlements once. #### Step 04 — The agent loop runs in the user's browser. Routing, tool decisions, plan execution, streaming: all on the client. No proxy hop before the LLM call. #### Step 05 — The server handles credentials and proxy duties. LLM calls, tool execution against your backend, audit. The session envelope is issued once and cached for five minutes. After that, the server isn't in the reasoning loop. #### Step 06 — Every turn is traceable. Each request stamped with a UID. Correlated across BigQuery (proxy + LLM events), Firestore (config + state), and the browser console (what the user saw). Reproduce any turn at any time. ### Teams shipping AI in weeks, not quarters. > "We went from a 9-month timeline to weeks — the Mindset stack let us build multiple use cases — coaching, compliance, and course discovery — without starting over each time." > — James Cranwell, Head of Product, 5app > "We built the Social Value Insight tool in 6 weeks. The feedback from our customers has been outstanding." > — Jacqui Bateson, Managing Director, HACT > "We built multiple specialist agents — each acting as a topic expert — and surfaced them into partner systems. We now have a $1M enterprise pipeline and 20% higher partner retention." > — Jessica Baker, Co-founder & CPO, AchieveUnite ### The questions a CTO asks before a demo. **Q: How does this compare to LangChain, LangGraph, or Vercel AI SDK?** Architecturally, Mindset AI is equivalent to LangGraph server-side: plans, tools, memory, citations, the full ReAct loop. The difference is where the execution lives, not what the framework can do. That difference shows up in three places: latency, page access, and iteration pace. **Q: WebSocket protocols solve part of this. What's different?** A persistent socket lets the server hold context between turns. Real improvement, but it ties you to one provider's protocol. Switch LLMs and you re-architect. Adopt it and you've traded one form of friction (context retransmission) for another (vendor lock-in). The fundamental issue is still there: the agent loop running on a stateless server. **Q: Won't bundle size kill our page load?** The orchestration runtime is ~45KB gzipped, loaded lazily after the host page is interactive. It does not block first paint. Smaller than most analytics scripts teams already ship. For hard bundle constraints (regulated environments, mobile-heavy audiences), we have deployment patterns that lazy-load the agent only when summoned. **Q: What happens if the user closes the tab mid-conversation?** Thread state is persisted to Firestore on every turn, indexed by thread. Re-opening loads the thread. The session envelope is cached for five minutes, so re-entry within that window is near-instant. Outside it, you re-issue the envelope. No state is lost. **Q: How do we get observability when the runtime is on the client?** Every turn is stamped with a requestUID that correlates across BigQuery (LLM and proxy events with cost, tokens, latency), Firestore (the config that was served, the message persisted), and the browser console (what the user saw). A support engineer can reproduce any turn at any time. **Q: Which browsers are supported?** Chrome 140+, Firefox 140+, Safari 17+, Edge 140+. Chrome, Firefox, and Edge auto-update, so most users are on a supported version by default. If your user base includes locked-down enterprise environments running older browsers, this is the first thing to confirm. **Q: Which LLM providers do you support?** Every major provider. Client-side orchestration is provider-agnostic. Switch providers without re-architecting. **Q: Does this lock us into Mindset AI?** No. Perpetual commercial licence on the Chat SDK, Headless SDK, and agent framework. Every widget exports as standard React. Speaks MCP for tools, OpenAI Chat Completions for models. One URL to adopt, one URL to leave. Updates stop if the relationship ends. Ownership doesn't. --- # HR & Recruitment Agents Webinar ## Ship a fleet of in-app HR & recruitment agents this quarter, without slowing your current roadmap. *Webinar for HR & recruitment software teams. Thursday 13 August 2026, 16:00 BST* Every HR and recruitment team needs agents inside their product. Most have a full roadmap and constrained resources. This 45-minute live build is the exact playbook for shipping a fleet quickly, with no hiring, nothing dropped from the plan, all from real implementation stories. ### What we'll cover in 45 minutes A live build, with a tight walkthrough around it. Here's the running order, one HR and recruitment software team's story, start to finish. #### Part 1: What they built How they chose what to build first and what to kill, and who got to build agents beyond core engineering. #### Part 2: Getting it live Getting agents into the product over MCP, widgets and APIs, and how candidate data and GDPR work in practice. #### Part 3: Making it better Knowing a shipped agent is working, and fixing it without slowing the roadmap or waiting on a release. ### The HR & Recruitment Agent Delivery Playbook - **Identify**: A triage sheet and build-or-kill criteria. - **Build**: Prototype-to-production checklist: MCP, widget and API gotchas, candidate-data compliance. - **Manage**: Who builds without an engineer, safely, alongside core engineering. - **Improve**: Know a shipped agent's getting better, fix it in minutes. - **Ship**: The 4-week delivery timeline. ### Presenters - Jack Houghton, CPO, Mindset AI - Ryan Soosaraj, VP of Professional Services, Mindset AI ### The questions worth asking before you register. **Q: Does it replace my framework?** No. Mindset plugs into it and makes it better. **Q: Do we keep our IP?** Yes. Open protocols, a perpetual licence, and every agent exports as React. **Q: Couldn't we just build this ourselves?** For a few agents, yes. The session is about what changes at fleet scale. **Q: Can non-engineers really build?** Yes. In a visual builder or over MCP, inside engineering's guardrails. **Q: What about candidate and employee data?** It's handled at the point of entry, multi-tenant, not per agent. **Q: Is this a pitch?** No. A meeting is bookable, but the session stands alone. **Q: Can't make the time?** Register anyway. We send the recording to everyone who signs up. --- # Campaign Pages ## Your SCORM content is locked away. Unlock it. *CONTENT TRANSFORMATION* Transform static SCORM packages into dynamic, searchable knowledge inside your AI agents — no migration required. ### The problem with SCORM today - **Content locked in proprietary formats**: Millions invested in training content sits trapped inside SCORM packages that AI systems cannot access or index. - **Not searchable or conversational**: Learners cannot ask questions about content buried in slide decks and click-through modules. - **Low completion, low engagement**: 20–30% completion rates because learners cannot find what they need when they need it. ### What Mindset AI SCORM ingestion gives you - **Bulk extraction**: Process entire SCORM libraries in minutes — no manual uploads or re-authoring. - **Semantic chunking**: Content is broken into meaningful, searchable elements that AI agents can retrieve contextually. - **Agent-ready knowledge**: Ingested content feeds directly into your AI agents for real-time conversational learning. - **Provider agnostic**: Works with SCORM 1.2 and 2004 packages from any authoring tool or LMS. ## Go from AI-powered to AI-first. *FREE RESOURCE BUNDLE* Three guides covering Ideation, Planning, and Implementation — everything your team needs to ship AI agents. ### What's in the bundle - **Guide 1 — Ideation — The top five AI features SaaS companies are shipping**: Forward-looking trends for e-learning platforms and AI-native products. Understand what leaders are building today. - **Guide 2 — Planning — How to integrate AI into your product — successfully and fast**: Strategies and best practices for product teams evaluating build vs. buy and managing AI integration risk. - **Guide 3 — Implementation — The guide for making your learning platform agentic**: A step-by-step framework for shipping AI-first learning experiences using the Mindset AI platform. ## Site search could be losing you sales. Fix it in days. *CRAFT CMS PLUGIN* Replace keyword search with an AI agent that understands what shoppers want, surfaces the right products, and guides them to purchase. ### Why site search matters more than you think - **Poor search = lost sales**: Up to 30% of site visitors use search. When it returns irrelevant results, they leave — and the revenue goes with them. - **Keyword search misses intent**: Shoppers describe what they want in natural language. Keyword matching cannot understand "something warm for hiking in winter". - **Months to build from scratch**: Building semantic search, intent classification, and product recommendations in-house takes quarters of engineering time. ### What the plugin does - **Semantic product search**: Understands natural language queries and returns products that match intent, not just keywords. - **Guided purchase flow**: An AI agent that asks clarifying questions and narrows results to the right product. - **Search analytics**: See what customers are searching for, what converts, and where search fails. - **Live in days**: Install from Craft's plugin store, connect your catalog, configure, and deploy. ## Interested in seeing how AI can benefit your company? *THE FASTEST WAY TO DEPLOY AI* Tell us about your goals and we will get back to you within one business day. ## Your engineers build AI. We handle how your customers experience it. *FOR ENGINEERING LEADERS* See how Mindset AI integrates with your existing architecture — ship conversational AI without rebuilding your stack. - **No infrastructure to own**: We run the agent platform and control plane — chat UIs, widgets, multi-tenant deployment — so your team stays focused on core product. - **API-first, integrates anywhere**: Connect your Bedrock, Snowflake, or internal APIs. Build MCP servers and we handle the presentation layer. - **Enterprise security and compliance**: ISO 27001 certified, SOC 2 compliant, tenant-isolated data, and full audit logging out of the box. - **Ship in weeks, not quarters**: Go live in days instead of 12 months. Your team builds the AI capabilities, we handle deployment and conversational UI. ## AI for Product Leaders | Mindset AI *FOR CPOs & PRODUCT LEADERS* ### Ship AI product features without a 6-month build. Give your product team the tools to build, test, and deploy AI agents — without consuming engineering cycles or disrupting your roadmap. **Problem:** Your customers want AI features. Your roadmap is full. Product teams are under pressure to ship AI capabilities, but building conversational interfaces, RAG pipelines, and multi-tenant agent infrastructure from scratch takes quarters of engineering time. Meanwhile, competitors are shipping faster. You need a way to let product managers prototype, test, and deploy AI features without waiting for the next sprint cycle. **Solution:** Mindset AI gives product teams direct access to ship AI. The Agent Management Studio lets product managers build working AI agents visually — define agent behavior, connect knowledge sources, configure widgets, and deploy to customers. Engineers stay focused on core platform while product ships AI features in weeks. Multi-tenant deployment, enterprise guardrails, and observability come built-in. #### Agent Management Studio Build and configure agents without code Product managers define agent behavior, connect knowledge sources, and configure conversation flows — then deploy to any tenant. - Visual agent builder with drag-and-drop workflow design - Knowledge source management with automatic RAG configuration - Per-tenant customization and white-label deployment #### Multi-Tenant Deployment Ship to thousands of customers from one control plane Deploy AI agents across your entire customer base with tenant-level isolation, per-tenant configuration, and centralized observability. - API-based tenant provisioning at scale - Isolated knowledge stores per tenant - Centralized analytics and performance monitoring #### Performance Testing Validate agent quality before shipping Run automated test suites against your agents to catch hallucinations, verify citations, and ensure quality before any customer sees them. - Automated test generation from your knowledge base - Citation accuracy and hallucination detection - Regression testing across agent updates #### Observability & Analytics Understand what users ask and how agents perform Built-in analytics track every conversation — question themes, user sentiment, resolution rates, and content gaps — so product teams can iterate with data. - Real-time conversation analytics and sentiment tracking - Content gap identification for knowledge base improvements - Usage patterns and adoption metrics per tenant ## Content Ingestion APIs | Mindset AI *CONTENT INGESTION* ### Ingest and sync the content your AI agents need — enterprise RAG-ready. Enterprise RAG APIs for app-wide content ingestion with tenant-level secure uploads across AWS, GCP, and Azure. PDFs, PPTs, MP4s, SCORM, and more. **Problem:** Your content is scattered. Your AI agents cannot reach it. Enterprise content lives across cloud storage, CMS platforms, learning management systems, and internal drives. Building the ingestion, processing, and retrieval pipeline to make this content available to AI agents is a major infrastructure project — one that delays your AI roadmap by months. **Solution:** Managed content ingestion that connects to everything. Mindset AI handles the entire content pipeline — ingestion from any source, intelligent processing and chunking, multi-tenant RAG architecture with strict tenant isolation, and optimized retrieval. Your team calls our APIs to sync content and your agents get instant, contextual access to the right knowledge. #### Ingestion APIs Connect cloud platforms and third-party systems RESTful APIs for rapid content ingestion from AWS S3, GCP Cloud Storage, Azure Blob, and third-party systems. Batch or real-time. - Cloud-native connectors for AWS, GCP, and Azure - Webhook-driven real-time sync - Batch processing for bulk content migration #### Intelligent Processing Automatic extraction, chunking, and indexing PDFs, PowerPoints, MP4s, SCORM packages, and documents are automatically processed — text extracted, images parsed, and content chunked by semantic boundaries. - Multi-format support: PDF, PPTX, DOCX, MP4, SCORM - Semantic chunking by topic and context boundaries - Image and table extraction with OCR #### Multi-Tenant RAG Isolated knowledge stores per customer Separate, secure content stores for each tenant with programmatic access control. Knowledge APIs ensure agents only access content they are authorized to use. - Strict tenant-level data isolation - Programmatic user and tenant access control - Per-tenant knowledge base management via API #### Optimized Retrieval Fast, accurate search with automatic re-indexing Lightning-fast retrieval through managed indexing and performance tuning. Content updates trigger automatic re-indexing with rollback support. - Sub-second retrieval across large knowledge bases - Automatic re-indexing on content updates - Rollback support for content version management --- # About ## We're on a mission to make technology feel more human Technology should amplify what makes us human — not replace it. We're building AI that works alongside people, turning complex tools into intuitive experiences and putting humans always in the loop. ## Technology is at its best when you forget it's there Most enterprise AI feels like it was built for machines, not people. We think that's backwards. Mindset AI exists to close the gap between what AI can do and what humans actually need — turning powerful technology into experiences that feel natural, intuitive, and genuinely useful. We build AI agents that work alongside your team as co-workers, not replacements. Agents that understand context, respect expertise, and make every person they work with more effective. The result isn't just better productivity — it's work that feels more human. Our customers deserve the very latest of what's possible, delivered in a way they can actually use. That means cutting-edge AI paired with thoughtful design, real safeguards, and humans always in the loop. ## Meet the team shaping the future of human-AI collaboration A group of passionate and brave technology experts set to build a future where AI takes half of your workload — but none of your purpose. Team members: Barrie Hadfield (CEO, Co-founder), Christine Davies (CXO, Co-founder), Jack Houghton (CPO, Co-founder), Wic Verhoef (CDO, Co-founder), Will Evans (CTO, Co-founder), Niamh Mulhall (Co-founder, Executive Director), Chris Whitehead (VP of Sales), Roxann Federici (Product Manager), Jonathan Cecil (Integrations Lead), Caroline Stewardson (Customer Success), Chetan Sharma (Delivery Team), Deepak Totlani (Delivery Team), Fred Zingg (Data & Integration), Hana Rasul (Customer Success), Maddie Lewis (Technical Enablement), Petya Stoyanova (Design Team), Rhea Arora (Product Team), Tarun Sogani (Delivery Team (QA Lead)), Deepak Khandelwal (Delivery Team), Rob Gorton (CFO), Ryan Soosayraj (VP of Professional Services), Upneet Singh (Delivery Team), Peter Worsley (Business Development Representative) ## What we stand for ### Dependable Our word is our bond. We are organized and reliable, take full responsibility, and trust each other. We commit to timelines and deliver work by the deadline and to a consistently high standard. ### Caring We act with kindness and empathy. We support others because everyone deserves respect — no exceptions. We care about our customers, colleagues, and everyone we engage with. ### Brave We are innovators. We're unafraid of new ideas and open to being proven wrong. We challenge each other and the status quo. We lean into change to push the boundaries of innovation. ### Passionate We have an overwhelming ambition to succeed. We believe in Mindset AI and love what we're building as a team. We're highly enthusiastic and bring our mastery, expertise, and unique flavor to work. ### Collaborative We understand the balance of going fast alone and reaching far together. We trust everyone to do their job and take initiative — and we're ready to support one another. ### No Bullshit We value different opinions. We're unafraid to share our opinion even if at times we're alone in our thinking. We give honest feedback and call out each others' bullshit — respectfully. ## Everything you need to know to thrive at Mindset AI MOS (Mindset Operating System) is how we run the company — a living system that defines our processes, culture, and ways of working. It evolves as we do. ### About MOS MOS defines how we work at Mindset AI. It documents our core processes, explains why they exist, and serves as a transparent resource for anyone working with or considering joining us. ### Why is MOS public? Transparency is a core value at Mindset AI. By openly documenting our ways of working, values, and priorities, we make a statement about who we are and how we operate. ### Who is MOS for? If you're reading this, MOS is for you. It's for anyone who works with or is considering working with us — whether you're a customer, partner, investor, employee, or prospective hire. It's here to help you understand who we are, why we exist, and what we believe in. ### Why does MOS exist? - **Vision:** Making technology feel more human — closing the gap between what AI can do and what people actually need. - **Mission:** Building a sustainable, creative, and diverse global company that positively impacts the lives of everyone we work with and those who rely on our platform. MOS is how we hold ourselves to that standard — the structure behind the culture, so the way we operate reflects the future we're building for our customers. ### Inspiration MOS is built on decades of experience, trial and error, and the wisdom of some of the most innovative thinkers of our time. We've been heavily inspired by critical thinkers and their work. - [Drive: The Surprising Truth About What Motivates Us](https://www.goodreads.com/book/show/6452796-drive) by Daniel Pink - [The Founders Survival Guide](https://www.goodreads.com/book/show/62621095-the-founder-s-survival-guide) by Rachel Turner - [Reinventing Organizations](https://www.goodreads.com/book/show/20787425-reinventing-organizations) by Frederic Laloux - [Traction: Get a Grip on Your Business](https://www.goodreads.com/book/show/18886376-traction) by Gino Wickman - [The Advantage](https://www.goodreads.com/book/show/12975375-the-advantage) by Patrick Lencioni - [The Four Steps to the Epiphany](https://www.goodreads.com/book/show/762542.The_Four_Steps_to_the_Epiphany) by Steve Blank - [The Goal: A Process of Ongoing Improvement](https://www.goodreads.com/book/show/113934.The_Goal) by Eliyahu Goldratt - [Holacracy: A Radical New Approach to Management](https://www.youtube.com/watch?v=tJxfJGo-vkI) by Brian Robertson - [Essentialism: The Disciplined Pursuit of Less](https://www.goodreads.com/book/show/18077875-essentialism) by Greg McKeown - [Slow Productivity: The Lost Art of Accomplishment Without Burnout](https://www.goodreads.com/book/show/197773418-slow-productivity) by Cal Newport If you'd like to read any of these books, we'll cover the cost (for employees). Simply purchase the book and submit an expense claim. --- # Customer Stories ## We start with your KPIs Every engagement begins by defining the results that matter to you — then we help you hit them. Real metrics, real timelines, real accountability. ## Future Talent Learning: 95.2% learner retention and 6 hours saved per coach each month **Read the full case study:** https://www.mindset.ai/library/case-studies/ftl **Industry:** Learning & Development / Apprenticeships **Challenge:** FTL needed to scale coaching support across 1,300+ apprenticeship learners without growing their coaching team proportionally. Coaches were spending significant time on routine learner questions, and FTL needed to prove an AI-assisted model quickly — within a three-month window. **Solution:** Built a per-course AI tutor on the Mindset stack that handles learner questions around the clock, surfacing only the conversations that genuinely need human coaching. Deployed across all active courses in under three months, integrating directly with existing content without a custom infrastructure build. **Results:** - **95.2% learner retention** — Across 1,300+ apprenticeship learners using the AI tutor - **6 hours saved** — Per coach each month, redirected to high-value conversations - **116k+ interactions** — Learner conversations handled by the AI tutor - **<3 months to deploy** — From kickoff to live across all active courses > "The Mindset platform let us scale coaching support without scaling our team. Our coaches are spending their time on high-value conversations, and our learners are staying engaged." > — Rubecca Ali, Head of Learning Technologies, Future Talent Learning ## Hult Ashridge: Shipped a virtual learning assistant across 33 client programs **Read the full case study:** https://www.mindset.ai/library/case-studies/hult-ashridge **Industry:** Corporate Education **Challenge:** Hult Ashridge needed to deliver a virtual learning assistant across 33 executive education programs simultaneously within the EF Education First context, with strict per-program data isolation. Building custom AI infrastructure would have diverted the engineering team from their core roadmap. **Solution:** Built a multi-tenant architecture on the Mindset stack with per-program agent isolation — 33 live agents, each scoped to its program's content and learners, sharing a single platform without shared state. The engineering team shipped without diverting from core product work. **Results:** - **33 live agents** — Deployed across executive education programs with full per-program isolation - **NPS 76 → 80** — Learner satisfaction improvement across participating programs - **Engineering on roadmap** — Core product team stayed focused throughout the entire rollout ## : Shifting a 2,600-course catalogue to product-led growth **Industry:** Maritime & Energy Safety Training **Challenge:** A maritime safety training marketplace had 2,600 courses structured for traditional browse-and-search, making it difficult for professionals to find the right certification pathway under time pressure. Discovery was passive and the provider wanted to shift to a conversation-driven, product-led growth model. **Solution:** Replaced catalogue browsing with an AI-powered discovery conversation on the Mindset stack, mapping learner needs to the right courses through 14 defined lifecycle actions. The new experience debuted at five customer conference sessions to accelerate commercial adoption. **Results:** - **2,600 courses** — Entire catalogue surfaced through conversational discovery - **14 lifecycle actions** — Defined to guide learners from first interaction to certification - **5 conference sessions** — Customer-facing events anchored around the new AI experience > "The shift from browse to conversation fundamentally changed how our customers discover training. It's the kind of experience we'd always wanted to build but never had the platform to do it quickly." > — , CTO, Maritime Training Marketplace ## : 20,000+ learners reached with AI coaching across web and mobile **Industry:** Learning & Development **Challenge:** A corporate training provider needed to extend AI coaching to enterprise clients at scale — across both web and mobile — without pushing delivery past the quarter. Reaching 200+ client organisations with a consistent AI experience required shipping fast without per-deployment custom work. **Solution:** Built AI coaching on the Mindset stack and shipped across web and mobile within the quarter, covering 20,000+ learners across 200+ client organisations. One platform, deployed at enterprise scale in weeks. **Results:** - **20,000+ learners** — Reached with AI coaching across web and mobile - **200+ client organisations** — Served from a single platform deployment - **Shipped within the quarter** — From kickoff to live, on time against the enterprise commitment > "We needed to move fast without compromising on quality across our client base. The Mindset stack gave us the infrastructure to do both." > — , Product Leadership, Corporate Training Provider ## : A 20-product agent architecture proven in 90 days **Industry:** Healthcare Workforce & Clinical Education **Challenge:** A healthcare workforce and clinical education platform needed to validate a complex multi-product AI architecture — 20 separate agents, each scoped to a different product domain — within 90 days. The architecture required sub-three-second retrieval across all agents without compromising domain isolation. **Solution:** Built 20 product-specific agents on the Mindset stack, each backed by a dedicated MCP server connected to the relevant API. Achieved 2–3 second retrieval across all agents and proved the full architecture in production within the 90-day window. **Results:** - **20-product architecture** — Proven in production, one dedicated agent per product domain - **One MCP server per API** — Clean isolation between every product domain's data and logic - **2–3 second retrieval** — Across all 20 agents in the live environment ## : Course completion up 3.7 points in 90 days with a per-course AI tutor **Industry:** Apprenticeships & Vocational Training **Challenge:** An apprenticeship training provider needed to improve completion rates across vocational programmes, where learner disengagement and knowledge gaps were going undetected until assessment. They needed visibility into where learners were struggling — at scale — without increasing tutor workload, and with zero tolerance for safeguarding risk. **Solution:** Deployed a per-course AI tutor on the Mindset stack that surfaces knowledge gaps through conversation, with safeguarding guardrails built in from day one. Within 90 days completion rates had improved measurably and the system had flagged genuine curriculum gaps at scale. **Results:** - **+3.7 points** — Course completion improvement in 90 days - **17.4% of questions** — Surfaced genuine knowledge gaps for curriculum review - **Zero safeguarding incidents** — Full safeguarding compliance maintained throughout > "We needed something that could scale across all our courses without adding tutor load — and that took safeguarding seriously from the start. This delivered both." > — , CTO, Apprenticeship Training Provider ## Major Learning Platform: 900% increase in course enrollments **Read the full case study:** https://www.mindset.ai/library/case-studies/online-learning-platform **Industry:** Online Education **Challenge:** Learners were overwhelmed by 5,000+ courses without clear starting points. Completion rates were low and certificate sales underperformed due to limited perceived value. **Solution:** Built a recommendation agent on the Mindset stack connected to the full course catalog. The agent asks clarifying questions, surfaces personalised recommendations in seconds, and was live without a custom build. **Results:** - **900% increase** — In course enrollments among users who interacted with the AI agent - **3.4x higher revenue** — Generated by AI agent users versus the overall test group ## BreakPoint: 30,000+ employees deployed across enterprise clients **Read the full case study:** https://www.mindset.ai/library/case-studies/breakpoint **Industry:** Leadership Development & Wellness **Challenge:** Traditional senior management training diluted as it cascaded through teams. Employees struggled with content accessibility and engagement. Break Point needed to scale teaching without relying on hierarchical information distribution. **Solution:** Built white-labelled web and mobile agents using the Mindset stack — unlimited agents, guided conversational search, integrated knowledge banks, and analytics. Deployed fast without building the agent infrastructure from scratch. **Results:** - **30,000+ employees** — Deployed nationally across enterprise clients including Shell, Manchester University, and Bentleys - **Strong NPS** — Among senior leadership teams - **Daily engagement** — Users actively asking questions every day > "The AI ability is dynamic, its individual. I can put a question in there to stimulate an answer thats really tailored to me, and thats what I find so powerful." > — Ollie Ollerton, Founder, Break Point Academy ## Training Industry: 1,400 signups in 3 months with an 11% conversion rate **Read the full case study:** https://www.mindset.ai/library/case-studies/training-industry **Industry:** Learning & Development **Challenge:** Training Industry had 200,000 monthly visitors browsing free content then leaving without converting. Thousands of articles, podcasts, and videos made discovery overwhelming, and legacy content was buried. **Solution:** Built Tia (Training Industry Assistant) on the Mindset stack — a white-labelled agent embedded in their membership portal in weeks, with real-time content syncing from WordPress, Wistia, and Podbean. **Results:** - **1,400 signups** — Free trial signups in the first three months - **11% conversion** — From free trial to paid membership at $30/month - **~150 new subscribers** — Generating tens of thousands in new recurring revenue ## Persimmon Life: 40% of growth managers' time saved in month one **Read the full case study:** https://www.mindset.ai/library/case-studies/persimmon-life **Industry:** At-Home Health & Beauty **Challenge:** Regional growth managers spent 60–80% of their time answering repetitive FAQs from 500 trained nurses. Resources were scattered across Google Drives and learning platforms, and healthcare compliance required accuracy guarantees. **Solution:** Built compliance-ready agents on the Mindset stack with accuracy controls and healthcare guardrails. Connected directly to existing Google Drive repositories with automatic content sync — no custom infrastructure required. **Results:** - **40% time saved** — For regional growth managers within the first month - **50%+ of target** — Achieved within the first month of launch - **Adoption doubling monthly** — Nurse usage growing every month > "The feeling from our providers is just, thanks. They're thankful they have a space for answers, where going for a marketing answer they go here, for a business answer you go there and inventory answer you go somewhere else. So the overwhelming word is thanks." > — Rachel, Regional Growth Manager, Persimmon Life ## 5app: Went from a 9-month timeline to weeks with the Mindset stack **Read the full case study:** https://www.mindset.ai/library/case-studies/5app **Industry:** E-learning SaaS **Challenge:** 5app needed to deliver meaningful AI innovation quickly while managing competing priorities and a complex multi-tenant SCORM environment. They lacked dedicated engineering resources for a custom AI build. **Solution:** Mindset AI provided Agent Chat SDK, multi-tenant APIs, SCORM ingestion pipeline, agentic workflows, and MCP server integrations. 5app built VeeCoach, an AI coach integrated into their product. **Results:** - **9-month timeline to weeks** — Built coaching, compliance, and course discovery without starting over each time - **<2 hrs/month** — Ongoing maintenance managed entirely by the product team - **Increased ACV** — New pricing model increased Average Contract Value > "We went from a 9-month timeline to weeks — the Mindset stack let us build multiple use cases — coaching, compliance, and course discovery — without starting over each time." > — James Cranwell, Head of Product, 5app ## HACT: £800k+ recurring revenue and 80 customers upsold **Read the full case study:** https://www.mindset.ai/library/case-studies/hact **Industry:** Social Impact & Housing **Challenge:** HACT faced pandemic-driven behavior shifts requiring digital transformation. Manual processes prevented scalability, and the organization had limited digitization experience while needing cost-effective solutions. **Solution:** Built the Social Value Insight tool on the Mindset stack in 6 weeks — automated data collection, simplified customer journeys, and a repeatable framework they could reuse for future tools without starting over. **Results:** - **£800k+ recurring revenue** — Moved to a scalable recurring revenue model - **80 customers upsold** — Converted into managed accounts within two months - **6-week setup** — Solution fully configured and launched - **5 days/month saved** — After two months of implementation > "We built the Social Value Insight tool in 6 weeks. The feedback from our customers has been outstanding." > — Jacqui Bateson, Managing Director, HACT ## AchieveUnite: $1M in enterprise opportunities with 20% partner retention boost **Read the full case study:** https://www.mindset.ai/library/case-studies/achieveunite **Industry:** Partner Training & Enablement **Challenge:** AchieveUnite needed products supporting a recurring revenue model. Their vast collection of books, assessments, and learning materials required on-demand delivery at the point of need for enterprise clients like HP, Lenovo, and Atlassian. **Solution:** Built multiple specialist agents on the Mindset stack — each acting as a topic expert across their IP, research, and content. Embedded into partner systems as a workplace tool without building agent infrastructure from scratch. **Results:** - **$1M in opportunities** — Enterprise pipeline generated - **40% faster onboarding** — Accelerated partner enablement speed - **20% retention boost** — Increase in partner retention rates > "We built multiple specialist agents — each acting as a topic expert — and surfaced them into partner systems. We now have a $1M enterprise pipeline and 20% higher partner retention." > — Jessica Baker, Co-founder & Chief Program Officer, AchieveUnite ## Awards & Certifications - Software Advice — Most Recommended 2025 - Software Advice — Best Customer Support 2025 - Capterra — Best Ease of Use 2025 - Capterra — Best Value 2025 - GetApp — Best Functionality & Features 2025 - ISO/IEC 27001:2022 Certified --- # Mindset Library > 79 posts covering AI agents, agent governance, MCP, and the future of software. For all posts in full, see [llms-blog.txt](https://www.mindset.ai/llms-blog.txt) ## How We Built The Agents On Our Website **Date:** 2026-05-07 | **Read time:** 8 min | **Category:** Blog The course performance agent on our website analyzes course data, diagnoses why learners are dropping off at specific lessons, and recommends concrete fixes. It's embedded in an admin dashboard. It knows which team the admin is looking at. It acts on the page. It starts working without being asked. This piece walks through every component that went into making it work, so you can see what it would take to build the same on your platform and decide whether the approach makes sense before you talk to anyone here. ‍ ## **The line between what's ours and what's yours** Some pieces are ours. The agent runtime, LLM orchestration, streaming, multi-turn conversation management, the configuration studio where product teams build and iterate on agents under engineering-defined governance, expert knowledge injection (skills), rich visual rendering of tool results in the chat (widgets), persistent memory across sessions and devices. Some pieces are yours. Your data, your UI, your business logic. Below is the complete list of what was built on the "your side" of that line. Each item is independent. You can build them in any order. Each one is useful on its own. Together they compound. If you only ever build A, you have a working diagnostic agent. Add B and the agent finds problems across your catalog. Add C, D, E, F, G in any order and each one extends what the agent can do without invalidating what came before. ‍ ## **A. Course detail endpoint** **What you build.** An API endpoint that, given a course ID, returns the course structure with per-lesson engagement data. Most learning platforms already have an endpoint that returns course details. The addition is engagement metrics alongside the content metadata. For each lesson (or slide, module, activity, whatever your platform calls its content units): position and title, content type (video, text, quiz, interactive), type-specific characteristics (video duration and chapters, text word count and images, quiz questions and pass threshold, interaction type for interactive content), how many learners completed this lesson, how many learners stopped here (the last lesson they engaged with before not returning), that number as a percentage, and median time spent. **What this unlocks.** The agent can diagnose why a specific course is underperforming. It connects engagement numbers to content characteristics. "Lesson 4 is a 6-minute video with no chapters and no interaction points. 38% of learners who reach it leave the course entirely. Median watch time is 2 minutes 22 seconds out of 6 minutes. Recommend splitting into three shorter segments with a knowledge check between each." The agent doesn't just say "lesson 4 has high drop-off." It reads the lesson metadata, identifies the structural problem, and names the fix. ‍ ## **B. Cross-course query endpoint** **What you build.** An endpoint that accepts filter, sort, and limit parameters and returns courses matching the criteria with headline performance metrics. The agent needs to answer questions like "which courses have completion rates below 50%?" or "what are the highest-rated courses?" or "show me mandatory courses for the warehouse team sorted by drop-off rate." The endpoint returns per course: title, type, whether it's mandatory, which teams or groups it's assigned to, enrollment and completion counts, completion rate, average rating, and the maximum single-lesson drop-off. This endpoint is wrapped in an MCP server: a lightweight service that makes the endpoint available as a tool the agent can call. We provide a specification and guide for building MCP servers. If you've never built one before, the wrapping itself is a few hours of work; the endpoint underneath is the part that's specific to you. The wrapping pattern is generic. The endpoint underneath can be a database query, a third-party API, a custom service, or a Snowflake warehouse query. The agent doesn't care; it calls the tool, the tool runs whatever's wrapped underneath. **What this unlocks.** The agent finds problems the admin didn't know to look for. Instead of browsing course by course, the admin asks a question and the agent queries across the entire catalog, identifies the outliers, and presents a prioritized list. Combined with A, this gives you the full triage-then-diagnosis flow: the agent finds the problem courses, then drills into each one to explain why. ‍ ## **C. Embed the agent** **What you build.** A `` element on your page and a one-line SDK initialization call. The agent renders in your page as a panel, a sidebar, or a full-screen experience. It runs on Mindset AI. The agent's appearance (colors, typography) is configured to match your product's look and feel. **What this unlocks.** Your admins (or learners, or managers) have an AI agent available inside your product. They can ask questions, get analysis, and receive recommendations without leaving the page they're working on. The agent has access to any MCP tools you've built (A, B) and any skills and widgets configured in the Agent Management Studio. This is the foundation for D, E, F, and G below. Those capabilities enhance the embedded agent, but each is optional and independent. ‍ ## **D. Page actions** **What you build.** Short JavaScript functions that wrap actions your admin UI already supports. If your admin interface has a "flag for review" button, a navigation link to the course editor, or a team filter dropdown, you can make those same actions available to the agent. Each one is 10 to 20 lines of JavaScript. No backend changes. No deployment. The function runs in the browser, in the admin's authenticated session, using whatever your UI already uses when the admin clicks the button manually. You control exactly which actions the agent can take, and you can vary them by page, by role, or by application state. A tool that isn't registered doesn't exist to the agent. **What this unlocks.** The agent moves from advisor to actor. After diagnosing a problem, it offers to flag the course for review (a badge appears in the UI), open the editor at the problem lesson (the editor opens), or filter the view to a different team (the filter changes). The admin sees things happen on the page. ‍ ## **E. Context** **What you build.** A few lines of JavaScript that tell the agent what the admin is currently looking at. Which page they're on, which filters are active, which item is selected. This updates live as the admin navigates. **What this unlocks.** The admin changes a filter. The agent picks up the new context and re-analyzes without being asked. Different team, different problems, different recommendations. No reloading, no new session. The admin never has to type "I'm looking at the warehouse team's courses"; the agent already knows. ‍ ## **F. Identity and multi-tenancy** **What you build.** The customer account ID and user role, passed as headers to your MCP endpoints. This enforces multi-tenancy: one agent serves all your customers, but each admin only sees their own data. The scoping is enforced by your endpoints, not by the AI. **What this unlocks.** One agent configuration serves your entire customer base. Data isolation is enforced server-side. You don't need a separate agent per customer. The same agent, the same skills, the same widgets, scoped to the right data by identity headers that your endpoints verify. ‍ ## **G. Programmatic triggers** **What you build.** One-line JavaScript calls that send a message to the agent from your application. Any event on your page can trigger the agent: a button click, a page load, a filter change, a record selection. The trigger can be visible in the chat or silent (the agent responds but the trigger message is hidden). **What this unlocks.** The page loads and the agent immediately starts analyzing the current view. The admin doesn't type anything. The admin clicks "analyze this course" and the agent begins a detailed lesson-by-lesson diagnosis. The admin changes the team filter and the agent automatically re-analyzes for the new team. The application orchestrates the agent, not just the user. Proactive, event-driven agent experiences become possible without the admin needing to know what to ask. ‍ ## **How they compose** | Combination | What the agent can do | |-------------|----------------------| | A alone | Diagnose why a specific course is failing | | A + B | Find problem courses across your catalog, then diagnose each one | | C + A + B | Agent embedded in your product, finding and diagnosing courses | | C + D + A + B | Agent acts on the page after diagnosis: flags courses, opens the editor | | C + E + A + B | Agent knows context, scopes analysis to the admin's current view | | C + F + A + B | Agent serves multiple customers from a single configuration | | C + G + A + B | Agent starts working on page load, re-analyzes on filter changes | This demo shows all seven working together. ‍ ## **Summary** | Capability | What you build | What the agent can do | |------------|---------------|----------------------| | A. Course detail | One enriched endpoint | Diagnose why a specific course is failing | | B. Cross-course query | One endpoint + MCP wrapper | Find problem courses across your catalog | | C. Embed the agent | Agent element + SDK init on your page | AI agent available inside your product | | D. Page actions | JavaScript functions wrapping existing UI actions | Act on the page: flag, navigate, filter | | E. Context | JavaScript on the host page | Agent knows what the admin is looking at | | F. Identity | Headers passed to your MCP endpoints | Multi-tenant: one agent, many customers | | G. Triggers | JavaScript calls to send messages to the agent | Proactive, event-driven agent behavior | Each is independently valuable. ‍ ## **What Mindset AI provides** Everything not listed above. The agent runtime, LLM orchestration, streaming, multi-turn conversation management, the configuration studio where product teams build and iterate on agents under engineering-defined governance, expert knowledge injection (skills), rich visual rendering of tool results in the chat (widgets), persistent memory across sessions and devices, and the infrastructure to run all of this at scale across your customer base. You build the domain pieces that only you can build. The platform handles the rest. ‍ ## **What this looks like for a different domain** The same primitives compose for any in-app agent. Our customer success demo uses A and B against account data instead of course data, adds skills (expert knowledge configured in the Agent Management Studio) for CS playbook interpretation, and uses page actions to mark up a portfolio chart and draft outreach messages. Same building blocks, different domain. The endpoints change, the page actions change, the knowledge changes. The pattern doesn't. ‍ ## **Where to go from here** If you want to see the same pattern applied to a use case closer to yours, the customer stories page has worked examples from teams who built on Mindset AI. If you want the technical specification for the MCP wrapper in B, it's in our docs. For worked MCP examples across healthcare, fintech, logistics, edtech, insurance and travel commerce, see the integration patterns doc. If you want to see what the studio looks like when product teams iterate on agents without engineering involvement, that's a separate demo we're happy to walk through. --- ## What Breaks When You Scale Agents From One To Many **Date:** 2026-05-07 | **Read time:** 7 min | **Category:** Blog Most agent frameworks focus on the model. The model is the easy part. The hard part is everything between the model and the user, and that's where teams break when they go from one agent to many. This piece walks through what we keep seeing break in production, and what we built into Mindset AI to handle it. Each section is a real failure mode that shows up in the second, fifth, or tenth agent. Not hypothetical edge cases. If any of them maps to where your team is right now, that's the point. ‍ ## **Every new agent combination means new code** You ship one agent. Then product asks for agents for onboarding, support, enterprise, and the power-user tier across three products. Each one needs different MCPs, knowledge sources, and permissions. Enterprise wants a different configuration than standard tier. A specific team wants different tools than another team. A reseller channel needs a white-labeled variant. A regional team wants agents scoped to their market. The model doesn't change between these combinations. What changes is which agent, which tools, which knowledge, which widgets, which permissions get assembled and delivered to each user. If every combination is a separate implementation, the number of combinations grows faster than the team. **What we did about it.** Agents, MCPs, knowledge contexts, skills, and widgets are independent building blocks. One API call assembles any combination per user and per tenant. ``` POST /api/v1/appuid/{APP_UID}/agentsessions { "agentUid": "support_agent_v2", "externalUserId": "user-x123456", "contextUids": ["product_docs", "enterprise_kb", "compliance_v3"], "tags": ["enterprise", "emea"] } ``` Standard tier, enterprise tier, per reseller, per regional package: no code change. The same building blocks that work for one customer work for ten thousand without the architecture changing. ‍ ## **Agent changes queue behind feature work** A system prompt update needs a PR. A widget change needs engineering. A new MCP tool connection needs a deploy. These are agent changes, not model changes, but they require the same engineering cycle as shipping a feature. Product can't iterate without engineering, and engineering can't prioritize agent improvements when they're building the next feature. Most companies' current stacks have no middle ground between "everything goes through engineering" and "open the codebase." Agent improvements are inherently iterative; you learn from user interactions and need to ship changes fast. When agent iteration requires the same deployment cycle as feature work, it queues behind it. The tradeoff is: gate everything through engineering and watch the backlog grow, or open the doors and risk prompt conflicts, broken tool connections, and inconsistent behavior across products. **What we did about it.** The Agent Management Studio lets product, CS, and other non-engineering teams build, prototype, and iterate on agents within engineering-defined governance, without touching the codebase. Engineering controls what ships, which tools are available, which models are used, what governance rules apply. Both sides work from the same place. System prompts, MCP connections, and widget updates publish without a deployment cycle. MCP credentials are injected server-side, access controlled centrally, auditable, reversible. ‍ ## **The agent can't see or do anything on the page** Most agent architectures treat the chat window as the only bridge between the agent and the application. The agent can answer questions. It can't see what the user is looking at, fill in a form field, highlight a row, navigate to a section, or take any action on the page. When a user says "help me with this," the agent doesn't know what "this" is. When it recommends an action, the user has to go do it manually. The gap between the agent's answer and the user's task completion is entirely on the user. Closing that gap with custom code means bespoke integrations per page, per component, per action. **What we did about it.** Five frontend primitives give agents controlled page access from browser JavaScript alone. Situational awareness: the host application tells the agent what the developer chooses to share about the user's current context. Page tools: the host registers JavaScript functions as tools the agent can call; the handler runs in the user's browser with full DOM access. Pass-through parameters: structured data flows directly to tools, bypassing the LLM. Agent session parameters: identity and routing flow to MCP servers as HTTP headers. sendMessage: the host application sends a message to the agent on behalf of the user, triggering a full response. The developer writes the handler and controls exactly what it does. A tool that isn't registered doesn't exist to the agent. It's a controlled capability grant, not autonomous access. ‍ ## **Something broke in the chain. Finding what is the hard part** A system prompt update, a new MCP server, a widget change: any of them can break something in the chain between the model and the user. There's no visibility into what changed, what it affected, or whether it's still working. Debugging means tracing through every component: prompt assembly, tool calls to MCP servers, widget rendering in the browser, page context, memory, with no single view. When something breaks, the discovery path is usually a customer reporting it. The fix path is usually hours of reconstruction. **What we did about it.** Mindset AI monitors the full last-mile chain from system prompt to tool call to rendered widget, with change tracking and failure detection before users report it. Changes can be tested, previewed, and rolled back without a deploy. ‍ ## **You don't know whether the agent is delivering value** The agent is live. Is it solving the use case it was built for? You don't know. The data you have is thumbs up, thumbs down, and reading user messages manually. None of that tells you whether users are completing tasks, where they're dropping off, which tools aren't being called, which tool descriptions don't match how users phrase their questions, or whether users are asking for capabilities that no tool covers. Backend evals test whether the model responds correctly to a prompt. They don't test the last mile from user question through tool calls to rendered result. The agent could be passing every eval and failing the actual use case. **What we did about it.** Mindset AI tests the full last-mile chain from user question to tool call to widget render to task completion, against real user interactions. Whether tools are calling correctly, whether tool descriptions match how users phrase their questions, whether widgets rendered, whether memory worked, and how agents perform against scenarios you define. Even when you know something's wrong, you don't know what to fix first. Mindset AI surfaces specific, ranked recommendations from real user interaction data: system prompt adjustments, tool description rewrites, missing knowledge contexts, widget fixes. Accept or reject. Agents self-improve from actual demand-side data. ‍ ## **More capable agents get slower on server-side frameworks** Every server-side framework adds a round-trip per tool call. As your agent gets more capable (more tools, multi-turn reasoning, richer context), latency increases linearly. With twenty or more tools active, that's six to eight seconds before anything renders in the browser. The tradeoff on every server-side framework is: make the agent more capable, or keep it fast. **What we did about it.** Mindset AI's framework runs tool orchestration in the browser. One to two seconds with twenty or more tools, no server round-trip per tool call. The framework is built on LangGraph and runs under perpetual license; if you want to fork it and run it independently of Mindset AI, you can. ‍ ## **One surface. Customers expect five** The agent works in your web app. Product wants it in the mobile app. A customer asks for Slack. Then Teams. Then WhatsApp. Each new surface is a separate integration, a separate deployment, a separate set of edge cases. When AI surfaces like Claude and ChatGPT mature, you'll need to be there too. **What we did about it.** A single agent setup deploys across web, mobile, and messaging platforms with the same widgets, tools, and behavior rendered per surface. Add a surface without rebuilding the agent. ‍ ## **Compliance applies at the frontend, not just the backend** GDPR, the EU AI Act, data residency: these requirements apply at the point where user data enters the agent system, which is the frontend. Questions, conversation history, context, personal information all pass through the frontend first. Retrofitting compliance into this layer is significantly harder and more expensive than building it in from the start. **What we did about it.** ISO 27001 certified, GDPR-ready deletion, data residency controls, EU AI Act alignment for agent transparency and auditability. Agent memory built to spec with compliance controls. Data stays where you set it, models you've already vetted. ‍ ## **What this adds up to** Most agent frameworks focus on the model. The model is the easy part. Every framework gets a working answer back in a chat window in a few hours. The hard part is everything that happens before that answer (which agent, which tools, which knowledge, which permissions, scoped to which user) and everything that happens after (whether the answer rendered correctly, whether the user actually completed the task, what to change next). That's the layer Mindset AI is built for. Composable building blocks. Engineering keeps architectural control while product teams iterate. Five frontend primitives that let agents see and act on the page. Last-mile observability and self-improvement from real user interactions. Browser-native execution that doesn't get slower as you add tools. One setup, every surface. Compliance from day one. Everything you build on Mindset AI is yours. Widgets export as standard React. The framework runs under a perpetual license. One URL to adopt, one URL to leave. --- ## How AI Is Changing Work (And Why Most Companies Are Behind) **Date:** 2026-02-26 | **Read time:** 8 min | **Category:** Blog Nearly every company claims to be investing in artificial intelligence, budgets are growing, pilot programmes are launching, and executives can't stop talking about AI workplace transformation in board meetings. But when you look at what's actually happening on the ground, the picture is very different. According to [Gallup's latest workforce data](https://www.gallup.com/workplace/701195/frequent-workplace-continued-rise.aspx), nearly half of U.S. workers, around 49%, say they never use AI in their roles. Not rarely, not occasionally, but never. Meanwhile, only 12% use it daily. That's a stark gap between the narrative companies are telling and the reality their employees are living. The AI-powered future of work that everyone keeps promising hasn't arrived for many people. And the longer companies wait to close that gap, the harder it becomes to catch up. ## The gap between talking about AI and actually using it If you listened only to earnings calls and press releases, you'd think every company had gone all-in on AI. And in one sense, that's true. [McKinsey's 2025 State of AI report](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai) found that 88% of organisations use AI in at least one function. That sounds impressive until you realise most haven't scaled beyond pilot programmes. They're experimenting, not operating. This is what researchers at [Slalom](https://www.slalom.com/us/en/insights/ai-research-outlook-2026) have called the "ambition execution gap," and their 2025 survey of 2,000 business and technology leaders confirms it's getting worse, not better. Companies are confident, and budgets are rising, but the day-to-day reality is uneven adoption, shallow integration, and fragmented governance. The same McKinsey research shows that only about 39% of organisations report bottom line EBIT impact from their AI investments at the enterprise level. That means the majority of companies spending significant money on AI can't yet point to measurable financial returns. They're running the demos and building the slide decks, yet the actual business outcomes are lagging well behind the investment. And it's not because the technology doesn't work. It's because most companies are approaching AI adoption the wrong way. ## Why pilot programmes aren't enough There's a pattern that keeps repeating across industries. A company launches an AI pilot in one department. Maybe it's marketing using a content generation tool, or finance testing an automated reporting workflow. The pilot goes well enough that leadership feels good about it. So they announce more pilots, maybe in customer support and operations. Six months later, the company has a dozen isolated experiments running simultaneously, each with its own tools, its own data practices, and its own definition of success. But none of them connect to a broader strategy, and none of them have meaningfully changed how the organisation works. According to [Deloitte's State of AI in the Enterprise report](https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html), only 34% of organisations are genuinely using AI to reimagine their business: creating new products, reinventing processes, or transforming business models. Another 30% are redesigning some key processes around AI. But more than a third are still using AI at a surface level, with little or no change to existing workflows. PwC puts it bluntly in their [2026 AI predictions](https://www.pwc.com/us/en/tech-effect/ai-analytics/ai-predictions.html): many companies take a ground-up approach to AI, crowdsourcing initiatives from individual teams and then trying to shape them into something resembling a strategy. The result is a collection of projects that don't match enterprise priorities, lack precision in execution, and seldom lead to actual transformation. You can create impressive adoption numbers this way, but you won't produce meaningful business outcomes. The companies getting real value from AI aren't running experiments. They're making strategic decisions about where AI fits into their core operations, and then building the organisational muscle to make it work. ## The leadership problem hiding in plain sight One of the most revealing findings from the Gallup data is the disconnect between leadership and everyone else. In Q4 2025, 44% of leaders reported using AI frequently at work, compared to just 23% of individual contributors. That gap has been [growing over time](https://www.gallup.com/workplace/701195/frequent-workplace-continued-rise.aspx), not closing. The people making decisions about AI strategy are drifting further from the people who'd actually use it day to day. This matters because AI adoption doesn't happen in a vacuum. Gallup's research shows that employees whose managers actively support AI use are more than twice as likely to use it frequently themselves. But only about 30% of workers say their manager provides that kind of support. The remaining 70% are either getting mixed signals or hearing nothing at all. At the organisational level, only 38% of employees say their company has actually integrated AI technology to improve productivity and quality. Another 41% say their organisation hasn't implemented AI tools at all, and 21% simply don't know. When a fifth of your workforce can't even say whether your company uses AI, you don't just have an adoption problem. You have a communication problem. The companies succeeding with AI share something in common: their senior leaders visibly champion AI initiatives, use the tools themselves, and give teams clear guidance on when and how to integrate them. McKinsey's research found that AI high performers are three times more likely to have this kind of senior leadership commitment than their peers. ## What's actually at stake Falling behind on AI doesn't just mean missing a trend. It means slower product development, higher operating costs, and a growing inability to compete with organisations that have already made AI part of their daily operations. PwC's [Global AI Jobs Barometer](https://www.pwc.com/gx/en/services/ai/ai-jobs-barometer.html) found that revenue growth in industries best positioned to adopt AI has nearly quadrupled since 2022. Wages in those same industries are rising twice as fast as in less AI-exposed sectors. The economic rewards are real. They're just going to the companies and workers who moved first. The [World Economic Forum's Future of Jobs Report 2025](https://www.weforum.org/publications/the-future-of-jobs-report-2025/) projects that 86% of employers expect AI and information processing technologies to significantly transform their business by 2030. That's a broad consensus that change is coming. But expecting transformation and being prepared for it are two very different things. A [Morgan Stanley survey](https://www.morganstanley.com/insights/articles/ai-adoption-accelerates-survey-find) of corporate executives across the U.S., Germany, Japan, and Australia found that AI adoption has already led to the elimination of 11% of jobs in surveyed organisations, with an additional 12% left unfilled. At the same time, 27% of employees had been retrained in the past 12 months. AI isn't waiting for companies to be ready. It's reshaping work now, and the companies that are actively reskilling their people are the ones setting the pace. ## What the companies getting it right are doing differently The organisations pulling ahead on AI adoption aren't necessarily the ones with the biggest budgets or the most advanced technology. They're the ones that have figured out a few things that most companies still haven't. The most important shift is treating AI as a strategy problem, not a technology problem. Instead of letting a hundred pilots bloom and hoping something sticks, they identify specific workflows where AI can generate measurable value and focus their resources there. PwC describes this as a top-down approach where senior leadership picks the spots for focused AI investments, looking for key processes where the payoff can be significant. They also invest in people, not just tools. [Deloitte's research](https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html) found that insufficient worker skills remain the biggest barrier to integrating AI into existing workflows. The companies making real progress are the ones treating education and reskilling as a core part of their AI strategy, not an afterthought. Worker access to AI rose by 50% in 2025, according to the same report, but access without training just creates confusion. And critically, they measure what matters. McKinsey's data shows that tracking defined KPIs for AI initiatives is the strongest predictor of bottom-line impact. Yet fewer than 20% of enterprises currently do this. Most companies are scaling AI faster than they're building the frameworks to understand whether it's actually working. Perhaps most importantly, they redesign workflows rather than layering AI on top of existing processes. McKinsey found that high-performing organisations are nearly three times more likely than their peers to completely redesign their workflows when implementing AI. They don't just give people new tools and expect results. They rethink how the work gets done in the first place. At Mindset AI, we see this pattern play out regularly. The companies that come to us aren't struggling because they lack ambition or budget. They're struggling because they've been spending engineering time on infrastructure that doesn't differentiate them, instead of focusing on the AI capabilities that actually make their product unique. The ones who shift that balance tend to move fast. ## The window is narrowing, but it hasn't closed yet The good news is that most companies are still early enough in their AI journey that catching up is possible. When only 12% of workers use AI daily and only 34% of organisations are truly reimagining their businesses around it, there's still time to get this right. But that window won't stay open much longer. Companies that invested early in strategy, training, and workflow redesign are now pulling away from the pack, generating measurable returns while others are still debating which tools to buy. Each quarter that passes without a clear AI strategy makes the next quarter harder. The AI-driven future of work isn't some distant concept anymore. It's happening in the organisations that decided to stop experimenting and start operating. The question isn't whether AI will transform how your company works. It's whether you'll be the one leading that transformation or scrambling to respond after your competitors already have. And if nearly half your workforce has never used AI at work, you already know which side of that divide you're on. --- ### Older Posts - [Choosing A Conversational AI Platform: The Technical Comparison Framework](https://www.mindset.ai/blogs/conversational-ai-platform-comparison-framework) (2026-02-24) - [The SaaS apocalypse: What Happens When Software Costs Collapse to Zero](https://www.mindset.ai/blogs/the-saas-apocalypse-what-happens-when-creation-costs-collapse-to-zero) (2026-02-20) - [Why Protocol Based AI Infrastructure Is Replacing Custom Integrations](https://www.mindset.ai/blogs/why-protocol-based-ai-infrastructure-is-replacing-custom-integrations) (2026-02-19) - [Why Building Your Own Agentic Frontend Takes Longer Than You Think (And Why That Matters Now)](https://www.mindset.ai/blogs/why-building-your-own-agentic-frontend-takes-longer-than-you-think-and-why-that-matters-now) (2026-02-17) - [SpaceX just merged with xAI to put a million data center satellites in space | In The Loop Episode 44](https://www.mindset.ai/blogs/in-the-loop-ep44-spacex-merged-with-xai) (2026-02-16) - [Why Your AI Agents Need Visual Intelligence (Not Just Text Responses)](https://www.mindset.ai/blogs/why-your-ai-agents-need-visual-intelligence-not-just-text-responses) (2026-02-12) - [The Agentic Frontend: What It Is And Why Every Product Needs One](https://www.mindset.ai/blogs/the-agentic-frontend-what-it-is-and-why-every-product-needs-one) (2026-02-10) - [11 AI Predictions That Will Shape Product Development In 2026](https://www.mindset.ai/blogs/11-ai-predictions-that-will-shape-product-development-in-2026) (2026-02-05) - [The Build vs Buy Question Every CTO Gets Wrong](https://www.mindset.ai/blogs/the-build-vs-buy-question-every-cto-gets-wrong) (2026-02-03) - [Everything You Need To Know About GPT-5.2 In 10 Minutes | In The Loop Episode 43](https://www.mindset.ai/blogs/in-the-loop-ep43-chatgpt-5-2-review) (2025-12-18) - [Code Red: "We're At A Critical Time For ChatGPT." | In The Loop Episode 42](https://www.mindset.ai/blogs/in-the-loop-ep42-code-red-at-openai) (2025-12-11) - [How To Decide What To Automate With AI For Your Team & Customers | In The Loop Episode 41](https://www.mindset.ai/blogs/in-the-loop-ep41-how-to-decide-what-to-automate-with-ai) (2025-12-04) - [Why ChatGPT Atlas Browser Won’t Take Down Google | In The Loop Episode 36](https://www.mindset.ai/blogs/in-the-loop-ep36-openai-atlas-browser) (2025-10-30) - [AI’s Just Made Robotics Interesting Again | In The Loop Episode 35](https://www.mindset.ai/blogs/in-the-loop-ep35-embodied-ai-world-models-and-robots) (2025-10-22) - [What Is AI Workslop & How To Fix It | In The Loop Episode 34](https://www.mindset.ai/blogs/in-the-loop-ep34-what-is-ai-workslop) (2025-10-15) - [Top Three Announcements From OpenAI DevDay 2025 | In The Loop Episode 33](https://www.mindset.ai/blogs/the-loop-ep33-openain-devday-2025-summarized) (2025-10-08) - [AI Agent Memory: Why Your AI Agents Keep Forgetting Everything (And How We Fixed It)](https://www.mindset.ai/blogs/ai-agent-memory-beta-release) (2025-09-29) - [Meta Ray-Ban Display Smart Glasses: Yay Or Nay? | In The Loop Episode 32](https://www.mindset.ai/blogs/in-the-loop-ep32-meta-ray-ban-display-glasses) (2025-09-24) - [What are AI Companions & Should They Be Legal? | In The Loop Episode 31](https://www.mindset.ai/blogs/in-the-loop-ep31-what-are-ai-companions) (2025-09-17) - [The Real Cost Of AGI—According To OpenAI | In The Loop Episode 30](https://www.mindset.ai/blogs/in-the-loop-ep30-the-real-cost-of-agi) (2025-09-09) - [Is The AI Bubble About To Burst? | In The Loop Episode 28](https://www.mindset.ai/blogs/in-the-loop-ep28-is-the-ai-bubble-about-to-burst) (2025-08-19) - [What’s Replacing SCORM—And Should SCORM Be Replaced Or “Just” Transformed?](https://www.mindset.ai/blogs/what-is-replacing-scorm) (2025-08-18) - [GPT-5 Review: Everything You Need To Know | In The Loop Episode 27](https://www.mindset.ai/blogs/in-the-loop-ep27-gpt-5-review-everything-you-need-to-know) (2025-08-12) - [What’s The Future Of SCORM With AI?](https://www.mindset.ai/blogs/what-is-the-future-of-scorm) (2025-08-11) - [Top Four AI Trends & Predictions Of Summer 2025 | In The Loop Episode 26](https://www.mindset.ai/blogs/in-the-loop-ep26-top-ai-trends-and-predictions-of-summer-2025) (2025-08-05) - [Why do people use SCORM?](https://www.mindset.ai/blogs/why-do-people-use-scorm) (2025-08-05) - [Why Is Corporate E-Learning So Bad & How To Fix It With AI? | In The Loop Episode 25](https://www.mindset.ai/blogs/in-the-loop-ep25-how-do-we-fix-edtech-with-ai) (2025-07-29) - [The New Playbook For Shipping AI Agents — Why Companies are Building on Mindset AI](https://www.mindset.ai/blogs/build-or-buy-a-conversational-ai-agent) (2025-07-21) - [What Is Context Engineering And Why Should You Care? | In The Loop Episode 23](https://www.mindset.ai/blogs/in-the-loop-ep23-what-is-context-engineering) (2025-07-09) - [How Do I Integrate AI Into My Product—Ideally By Yesterday](https://www.mindset.ai/blogs/how-do-you-integrate-ai-into-a-product) (2025-07-07) - [What Jobs Will AI Create—And Do The Luddites Have A Point? | In The Loop Episode 22](https://www.mindset.ai/blogs/in-the-loop-ep22-what-jobs-will-ai-create) (2025-07-02) - [New Release: Mindset AI SDK 2.4 - Fonts Customization](https://www.mindset.ai/blogs/sdk-2-4-agent-fonts-customization) (2025-07-01) - [How Enterprise CIOs Build & Buy Gen AI In 2025 | In The Loop Episode 21](https://www.mindset.ai/blogs/in-the-loop-ep21-how-enterprise-cios-build-and-buy-ai-in-2025) (2025-06-25) - [New Release: Mindset AI SDK 2.2 Multi-Tenancy Agents & Session Control](https://www.mindset.ai/blogs/mindset-ai-sdk-2-2) (2025-06-24) - [Three Reasons Why Apple Is Cooked | In The Loop Episode 20](https://www.mindset.ai/blogs/in-the-loop-ep20-apple-wwdc-and-the-illusion-of-thinking) (2025-06-18) - [Mary Meeker AI Trends 2025: Three Reasons Why AI Is Different From Any Other Tech In History | In The Loop Episode 19](https://www.mindset.ai/blogs/in-the-loop-ep19-mary-meeker-ai-trends-report-highlights) (2025-06-11) - [New Release: Mindset AI SDK 2.1 Theme Customization](https://www.mindset.ai/blogs/mindset-ai-sdk-2-1) (2025-06-10) - [What Is The Difference Between A2A And MCP? [With Videos]](https://www.mindset.ai/blogs/what-is-the-difference-between-a2a-and-mcp) (2025-06-09) - [What Happens To Entry-Level Jobs In The AI Era? | In The Loop Episode 18](https://www.mindset.ai/blogs/in-the-loop-ep18-will-ai-replace-entry-level-jobs) (2025-06-04) - [Mindset AI Appoints Pip White as Non-Executive Director](https://www.mindset.ai/blogs/mindset-ai-appoints-pip-white-as-non-executive-director) (2025-06-02) - [Google I/O & Microsoft Build In 10 Minutes: What We Learned From The Two Biggest AI Conferences | In The Loop Episode 17](https://www.mindset.ai/blogs/in-the-loop-ep17-microsoft-build-and-google-io-2025-summary) (2025-05-28) - [New Release: Mindset AI SDK 2.0](https://www.mindset.ai/blogs/mindset-ai-sdk-2-0) (2025-05-26) - [The Top Five AI Features SaaS Companies Are Shipping In 2025 (And Why They Work) | In The Loop Episode 16](https://www.mindset.ai/blogs/in-the-loop-ep16-top-five-ai-features-saas-companies-are-shipping-in-2025) (2025-05-21) - [Google, OpenAI, Meta, Anthropic & The Three Battles To Own All AI | In The Loop Episode 15](https://www.mindset.ai/blogs/in-the-loop-ep15-the-three-battles-to-own-all-ai) (2025-05-14) - [Should Conversational AI Agents Get Priority On Your E-Learning Platform’s Roadmap?](https://www.mindset.ai/blogs/should-conversational-ai-agents-be-on-your-roadmap) (2025-05-12) - [The Real State Of AI Adoption In 2025: What's AI Actually Used For? | In The Loop Episode 14](https://www.mindset.ai/blogs/in-the-loop-ep14-what-is-the-most-common-use-of-ai-today) (2025-05-06) - [In The Loop Episode 13 | Cluely: The AI App That Made Cheating Viral—And Maybe Acceptable?](https://www.mindset.ai/blogs/in-the-loop-ep13-cluely-the-ai-app-that-made-cheating-viral) (2025-04-30) - [In The Loop Episode 12 | Google Agent2Agent (A2A): The Future Of AI Agent Protocols Or A Flop?](https://www.mindset.ai/blogs/in-the-loop-ep12-what-is-a2a-protocol) (2025-04-23) - [How To Turn Your E-Learning Business Into An AI Coaching Solution](https://www.mindset.ai/blogs/how-to-turn-your-elearning-business-into-an-ai-coaching-solution) (2025-04-22) - [In The Loop Episode 11 | Shopify Memo: No Humans Hired Without AI Approval—Tobias Lütke's Vision](https://www.mindset.ai/blogs/in-the-loop-ep11-shopify-memo-no-humans-hired-without-ai-approval) (2025-04-15) - [Mindset AI Raises £4.3 Million To Meet Growing Demand For Embedded AI Agents For SaaS Businesses](https://www.mindset.ai/blogs/ai-agent-fundraising-announcement) (2025-04-14) - [In The Loop Episode 10 | Does ChatGPT's Viral Image Generator & The Ghibli Craze Spell The End Of Art & Creativity?](https://www.mindset.ai/blogs/in-the-loop-ep10-does-chatgpt-image-generator-spell-the-end-of-art) (2025-04-07) - [How To Monetize Your AI Agents: A Product Leader's Guide To Revenue Generation In EdTech](https://www.mindset.ai/blogs/how-to-monetize-an-ai-agent) (2025-04-07) - [In The Loop Episode 9 | Apple’s AI Crisis Exposed: Is It Having A Nokia Moment?](https://www.mindset.ai/blogs/in-the-loop-ep9-is-apple-falling-behind-on-ai) (2025-04-01) - [In The Loop Episode 8 | Model Context Protocol (MCP): The Newest AI Buzzword Explained](https://www.mindset.ai/blogs/in-the-loop-ep8-what-is-mcp-for-ai-agents) (2025-03-26) - [Agentic AI 101: Everything You Ever Wanted To Know About AI Agents But Never Dared Ask](https://www.mindset.ai/blogs/what-is-agentic-ai) (2025-03-26) - [In The Loop Episode 7 | Vibe Coding: Will Developers Be Out Of A Job In Six Months? Dario Amodei’s Take](https://www.mindset.ai/blogs/in-the-loop-ep7-what-is-vibe-coding) (2025-03-19) - [When To Use Agentic RAG—And What Is It Anyway?](https://www.mindset.ai/blogs/when-to-use-agentic-rag) (2025-03-18) - [In The Loop Episode 6 | Multi-Agent Systems: The Next Big Shift In AI—Yet People Have No Clue About Them](https://www.mindset.ai/blogs/in-the-loop-ep6-multi-agent-systems-the-next-big-shift-in-ai) (2025-03-14) - [In The Loop Episode 1 | DeepSeek’s AI Breakthrough: Hype or Game-Changer? A No-Nonsense Breakdown](https://www.mindset.ai/blogs/in-the-loop-ep1-whats-so-great-about-deepseek) (2025-03-11) - [In The Loop Episode 5 | The Rise Of Vertical AI Agents: Why SaaS Companies Should Be Worried](https://www.mindset.ai/blogs/in-the-loop-ep5-the-rise-of-vertical-ai-agents) (2025-03-05) - [In The Loop Episode 4 | Why Microsoft's CEO Thinks Everyone's Wrong About AI Agents & AGI](https://www.mindset.ai/blogs/in-the-loop-ep4-breaking-down-satya-nadellas-vision-for-agi) (2025-02-26) - [AI Expert Interview: The Benefits And Drawbacks Of Agentic AI](https://www.mindset.ai/blogs/the-benefits-of-ai-agents) (2025-02-25) - [In The Loop Episode 3 | The Real AI Challenge: Designing Human-Agent Interfaces That Work](https://www.mindset.ai/blogs/in-the-loop-ep3-designing-human-agent-interfaces-that-work) (2025-02-19) - [AI Agents vs. Everything AI: All The Definitions You'll Ever Need](https://www.mindset.ai/blogs/ai-agents-vs-other-ai-paradigms) (2025-02-18) - [In The Loop Episode 2 | The Future of AI Agents: What’s Real, What’s Hype & What’s Next](https://www.mindset.ai/blogs/in-the-loop-ep2-what-are-ai-agents) (2025-02-12) - [When Did AI Agents Become A Thing? The History & Evolution Of Agentic AI](https://www.mindset.ai/blogs/how-have-ai-agents-evolved-over-time) (2025-02-11) - [What Is The Future Of Agentic AI: Eight Predictions From A CPO](https://www.mindset.ai/blogs/what-is-the-future-of-agentic-ai) (2025-01-31) - [How To Use AI Agents To Fix Broken Search In Learning Platforms](https://www.mindset.ai/blogs/how-to-use-ai-agents-to-fix-broken-search-in-learning-platforms) (2025-01-14) - [The OpenAI Announcement Will Transform The Way Mindset AI Agents Engage With Users And Knowledge](https://www.mindset.ai/blogs/the-openai-announcement-will-transform-the-way-mindset-ai-agents-engage-with-users-and-knowledge) (2024-05-16) - [Why Learners Choose Google Over Your Learning Platform And How AI Can Change That](https://www.mindset.ai/blogs/why-learners-choose-google-over-your-learning-platform-and-how-ai-can-change-that) (2024-01-18) - [How AI can make your video content more interactive and engaging](https://www.mindset.ai/blogs/how-ai-can-make-your-video-content-more-interactive-and-engaging) (2024-01-10) - [ChatGPT Intellectual Property Issues: How To Protect IP](https://www.mindset.ai/blogs/chatgpt-ai-chatbot-risks-to-your-intellectual-property-ip-and-how-to-protect-it) (2023-11-28) - [The Future Of HR: AI-Powered Knowledge Assistants](https://www.mindset.ai/blogs/the-future-of-hr-ai-powered-knowledge-assistants) (2023-11-15) - [How AI Can Support Self-Guided Employee Onboarding And Reduce Ramp Times](https://www.mindset.ai/blogs/how-ai-can-support-self-guided-employee-onboarding) (2023-09-04) - [The Future Of Knowledge Management: How AI Will Change The Way We Manage Knowledge](https://www.mindset.ai/blogs/how-ai-will-change-the-way-we-manage-knowledge) (2023-08-07) --- # Developer Tools ## Overview **Section:** Get started / undefined Mindset AI makes it easy for you to build, manage and deploy AI agents inside your SaaS applications. ## The three pillars **Agent Management Studio (AMS).** A no-code interface for building, testing, customizing and monitoring agents. **SDKs.** Embedded UI components that surface agents you build in the AMS for your customers and users. You can also use our headless SDK to put agents into your existing UI. We also enable you to integrate different parts of the AMS into your app, so your users can build, manage and improve agents too. **APIs.** Tools for integrating agents throughout your experience to control permissions, so they can see your existing web-page UI, act on pages and more. --- ## Two ways to put an agent in your product **Section:** Get started / undefined You can embed a Mindset AI agent into your own application in one of two ways. Both use the same backend handshake and the same agent. What differs is how much of the interface you write. ## The two paths The drop-in element. One script tag and one HTML tag, and a working chat interface appears on your page. We render it, we style it, we handle the streaming and the state. You write no interface code. The headless client. The core functionality is available without our UI: same conversation functionality with no user interface at all. You get events as the agent thinks and replies, and commands to drive it. Every pixel is yours. **Good to know:** The visible `` tag supports exactly the same set of methods, properties and events that the UI-less variant does. | | Drop-in element | Headless client | |---|---|---| | What you write | Two tags and a session function | Your entire interface | | What we render | A full chat panel | Nothing | | How you load it | Script tag | ES module import by URL | | Time to first message | Minutes | Hours to days, depending on your design | | Styling | CSS custom properties we expose | However you like | | Good for | Support widgets, in-app assistants, anything where our chat UI is fine | Products where the agent has to look like it was always part of your app | ### Which to pick Start with the drop-in element. It's faster to get running, and it's the same underlying contract, so moving to the headless client later isn't a rewrite of your backend. Choose the headless client when the conversation needs to live inside an interface you already have, when the agent's output drives something other than a chat transcript, or when your design system won't tolerate a component you don't control. You can also mix them. The element is a thin shell over the same commands the headless client exposes, with no private channel to the runtime, so there's nothing the element can do that you can't. If you do want to mix them, make sure you include our full SDK build that includes our UI, even if you don't want to render it on the page. ## How trust works The important thing to understand before you write any code: your organization's API key never reaches a browser. ``` Your backend Mindset AI ------------ ---------- Holds the org API key ----> Creates a session (server to server) <---- Returns an opaque credential | | hands the credential to the browser v Your page --------- Holds the session credential ----> Runs the agent (one agent, one user) <---- Streams the reply ``` Your backend makes one server-to-server call with your organization's API key and gets back a credential. That credential is scoped to one agent, one user, one organization and one Environment. It's the only thing that goes to the browser. The API key is admin-grade. The session credential is not: it can run one agent and nothing else, and it's refused on every management surface regardless of who the user is. Even an org admin's session credential can only run the one agent it was made for. The SDK keeps the session alive by itself. You don't write refresh logic, and there's no renewal endpoint for you to call. ## What runs where The agent runtime runs in the browser, not on our servers. Your page loads it as part of the SDK bundle, and it orchestrates the conversation from there, calling back to us for inference and tool execution. That's what makes page tools possible: the agent can call a function you define, running in your page, with your application's context. Your knowledge, your agent configuration and your conversation history stay on the Mindset AI platform. ## What you don't have to build **Session renewal.** The SDK handles it. When it can't, it asks your backend for a new session, which is the same call you already wrote. **Streaming.** Replies arrive token by token, and both paths hand you that stream. **Conversation resume.** The element gives you a conversation ID after a user's first completed turn. Store it against that user, hand it back on their next visit, and they pick up where they left off. **Style isolation.** The drop-in element renders in a shadow root with its own compiled stylesheet, so your CSS and ours can't interfere with each other. ## What isn't here yet Worth knowing before you plan around something that doesn't exist. **No conversation management API.** You can resume a specific conversation by handing back its ID, but there's no way to list a user's conversations, rename one, delete one, or switch between them programmatically. Resume is configuration rather than a command. **No self-hosted SDK bundle.** The SDK is served from your Mindset AI host and loaded by script tag or URL import. There's nothing to install, and you always get the current build. **No published rate limits.** We don't publish limits for the embed surface today. The SDK is additive-only, meaning new events and new fields appear but existing ones don't change shape or disappear. Build consumers that ignore anything they don't recognize and you'll be fine across updates. ## Where to go next - "Put an agent on your page" is the ten-minute quickstart. Start here. - "Create a session for your users" covers the backend call in full. - "The mindset-agent element" is the reference for the drop-in path. - "The UI-less SDK" is the reference for building your own interface. --- ## Put an agent on your page **Section:** Get started / Build it Two lines of HTML and one endpoint on your backend. This walks through both, and by the end you'll have a working agent your users can talk to. Budget about ten minutes if you already have an agent set up. ## What you need first **An agent, published.** Build it in the Mindset AI console. You'll need its handle, which you can copy from the agent's Embed tab. **An organization API key.** Ask your Mindset AI administrator for one. It's admin-grade and lives on your server only. **Your host address, org slug and Environment slug.** All three are on that same Embed tab, pre-filled for your organization, so you can copy them rather than assemble them. ## Step 1: add a session endpoint to your backend Your page can't call us directly, because doing so would mean putting your organization's API key in a browser. Instead your backend makes one call and hands the result to the browser. Add an endpoint that does this. Express shown here, but any stack works. ```javascript app.get("/api/session", async (req, res) => { const user = req.user; // however your app knows who's signed in const r = await fetch( `https://${MINDSET_HOST}/api/v1/orgs/${ORG_SLUG}/envs/${ENV_SLUG}/agent-sessions`, { method: "POST", headers: { "x-api-key": process.env.ORG_API_KEY, "content-type": "application/json", }, body: JSON.stringify({ user: { email: user.email }, agent: "support-bot", createUserIfNeeded: true, }), }, ); if (!r.ok) { const detail = await r.text(); console.error("Mindset AI session create failed", r.status, detail); return res.status(502).json({ error: "Could not start a session" }); } const { session } = await r.json(); res.set("cache-control", "no-store").json({ session }); }); ``` Three things to get right here. Your organization API key stays on the server. It's never sent to the browser, never in a response body, never in a log line. `createUserIfNeeded` set to `true` means a user we haven't seen before is created on their first visit. Leave it out and an unknown user gets a 404, which is the most common first-integration surprise. The `no-store` cache header stops a proxy handing one user's session to another. Return only the `session` value. It's opaque, so pass it along exactly as you received it rather than reading or reshaping it. For the full request and response detail, including how to use your own user IDs instead of email, see "Create a session for your users". ## Step 2: put the agent on your page Load the script and write the tag. ```html ``` That's the whole frontend integration. The script self-registers the `mindset-agent` tag, and `configure()` tells it how to get a session. Notice that `getSession` is a function, not a string. The SDK calls it when it needs a session, which is what lets it get a fresh one after a page reload without you writing any refresh logic. The agent renders inside a shadow root with its own styles, so it won't inherit your page's CSS and it won't leak styles into your page. ## Step 3: check it works Load the page. You should see a chat panel, and you should be able to send a message and get a reply. If nothing appears, open your browser console. The element reports configuration and transport problems as `mindset:error` events with a stable code you can branch on, and it logs a load failure rather than throwing. ```javascript document.querySelector("mindset-agent").addEventListener("mindset:error", (e) => { console.error("Mindset AI error:", e.detail.code, e.detail.message); }); ``` ## Using React React 19 renders custom elements natively, so there's no wrapper package to install. Same script tag, same element. ```jsx import { useEffect, useRef } from "react"; function Agent({ agent }) { const ref = useRef(null); useEffect(() => { ref.current?.configure({ getSession: async () => { const r = await fetch("/api/session"); if (!r.ok) throw new Error("Could not start a session"); const { session } = await r.json(); return session; }, }); }, []); return ; } ``` Load the script once, in your `index.html` or wherever you keep third-party tags, rather than inside the component. If you're on TypeScript, you'll need to declare the element in `JSX.IntrinsicElements` once, since a custom element isn't there by default. ## If you want to build your own interface Everything above gives you our chat UI. If you'd rather render your own, there's a headless client that gives you the same conversation with no UI at all. ```javascript const { createAgentConversation } = await import( "https://YOUR-MINDSET-HOST/sdk/mindset-agent-uiless.js" ); const chat = createAgentConversation({ agent: "support-bot", getSession: async () => (await (await fetch("/api/session")).json()).session, }); chat.on((event) => { if (event.type === "text_delta") appendToYourUI(event.content); }); chat.send("Hello"); ``` Your backend endpoint from step 1 is unchanged. Only the frontend differs. ## Before you ship **There's no npm package.** The SDK is served from your Mindset AI host and loaded by script tag or by URL import. There's nothing to install, and you always get the current build. **Only load one Mindset AI SDK per page.** A custom element tag can only be claimed once per document. Loading our bundle twice is harmless, but loading it alongside an older Mindset AI embed will collide, and the SDK will tell you so by name. **Conversations start fresh on every page load unless you tell them not to.** If you want a user to pick up where they left off, the element hands you a conversation ID after their first completed turn, and you hand it back on their next visit. See "The mindset-agent element" for the round trip. ## Troubleshooting **Nothing renders and the console says the script failed to load.** Check the host in your script tag. It should be the same host as the Embed tab shows. **`missing-agent`.** The element was configured without an agent. Check the `agent` attribute is set and matches a handle in your organization. **`invalid-auth`.** `configure()` didn't get a `getSession` function. A common cause is passing the session string itself rather than a function that returns one. **`turn-failed`.** The turn couldn't run. Usually the session expired or your session endpoint returned something unexpected. Check your backend logs first. **A 404 from your session endpoint during setup.** Four different things return an identical 404: a wrong org slug, a wrong Environment slug, an agent handle that isn't in your organization, and a key without the right scopes. The response won't tell you which, so check all four against the Embed tab. ## Next steps - "Create a session for your users" covers the mint call properly, including using your own user IDs. - "The mindset-agent element" is the full reference: attributes, methods, events, theming. - "The UI-less SDK" covers building your own interface. --- ## Create a session for your users **Section:** Get started / Build it Before an agent can run on your page, your backend makes one call to us. That call is the whole server-side integration. Everything after it happens in the browser. Your backend holds your organization's API key and never lets it near a browser. The browser holds a short-lived credential, scoped to one agent and one user, that your backend hands it. ## The shape of it 1. Your user loads a page that has an agent on it. 2. Your frontend asks your backend for a session. 3. Your backend calls us with your org API key and gets back an opaque credential. 4. Your backend passes that credential to the browser. 5. The SDK runs the agent with it. There's no step 6. You don't parse the credential and you don't store it. ## Before you start You'll need an organization API key. These are issued in the Mindset AI console by an org admin, so if you don't have one, ask your Mindset AI administrator to create one for you. Your key is bound to one Environment. A key for your test Environment can't create sessions in production, which is the behavior you want but does mean you'll need one key per Environment you're integrating. ## Make the call ``` POST /api/v1/orgs/{orgSlug}/envs/{envSlug}/agent-sessions ``` Your organization and Environment are both in the path. Send your key on the `x-api-key` header, not on `Authorization`. This route is server-to-server only and is never CORS-enabled, so it can't be called from a browser even by accident. ### Request body | Field | Type | Required | Notes | |---|---|---|---| | `user.email` | string | one of | Must be email-shaped. Creates a shared identity the same person can also log into the web app with | | `user.externalId` | string | one of | An opaque ID of your own. Letters, numbers, underscores and hyphens only, 100 characters or fewer. Embed-only, never web-loginable | | `agent` | string | yes | The agent's handle, or its UUID. Copy it from the agent's Embed tab in the console | | `createUserIfNeeded` | boolean | no | When true, a user we haven't seen before is created. Defaults to false, which means an unknown user is rejected | | `attribution` | object | no | Your own tags for this session. Up to 16 keys, keys up to 64 characters, values up to 256 characters. See below | Send exactly one of `user.email` or `user.externalId`. Sending both, or neither, is rejected. The body is strict. Any field not in this table is rejected rather than ignored, so a typo surfaces immediately instead of silently doing nothing. ### Example ```bash curl -sS -X POST \ "https:///api/v1/orgs/{orgSlug}/envs/{envSlug}/agent-sessions" \ -H "x-api-key: $ORG_API_KEY" \ -H "content-type: application/json" \ -d '{ "user": { "email": "alice@acme.com" }, "agent": "support-bot", "createUserIfNeeded": true, "attribution": { "plan": "enterprise", "region": "emea" } }' ``` ## What you get back A JSON object. Pass it to the browser exactly as you received it. The credential is opaque on purpose. Its internal shape isn't part of the contract and will change, and the SDK always gets back the structure it expects. Reading its fields, rebuilding it, or storing it in a database are all things that will break on you. One call does three things: it resolves or creates the user in your organization, confirms that user can reach that agent in that Environment, and returns the credential. Holding your org key and making this call is the authorization. The credential works for the named agent regardless of that agent's publication settings, as long as the agent has a published version to embed. ## Choosing an identity Use `email` when the person using your embedded agent is the same person who might log into Mindset AI directly. The same email is one identity across both, once they've proved they own the mailbox. Use `externalId` when your users exist only inside your application, or when you'd rather we didn't hold their email at all. It's opaque, it's never web-loginable, and it never merges with an email identity. We check the shape of an email, not whether it's real, so an address that can't receive mail will still work. What we won't accept is a value with no `@` in it, because that could later collide with an external ID and make one person look like two. Pick one and stay with it. Switching a user from one form to the other gives you two separate users with two separate conversation histories. ## Handing the credential to the browser Create a credential when your user opens the page, and use it straight away. A credential is meant to be consumed as soon as it's generated. It's single use, and it's scoped to one user and one agent. Keep it in memory. Don't put it in local storage, don't put it in a cookie, and don't render it into server-side HTML that gets cached. Your frontend gives the SDK a function that returns the credential, rather than the credential itself. That's what lets the SDK ask again when it needs to. Two agents on one page means two calls and two independent sessions. ## Keeping the session alive You don't have to do anything here. The SDK handles it, and there's no renewal endpoint for you to call. A session stays alive while the agent stays in place on the page. When it can't continue, because the tab closed or the page reloaded, the SDK asks your integration for a session again. Your backend does what it did the first time: create a new one and hand it over. Your endpoint stays a single create. ## Ending someone's access There are two ways to stop a user reaching an agent, and both are done in the Mindset AI console rather than through an API: - Remove the agent, which stops everyone reaching it. - Disable the user, which stops that person reaching anything. There's nothing more granular than this today, and there's no per-session revocation call. ## What the credential can and can't do The credential is admitted on the run surface only: inference, tool execution, and run events. It's refused on every management and administrative surface. That holds regardless of the user's role in your organization. An org admin's session credential still only runs one agent. There's deliberately no way to widen it. There's no restriction on where the SDK and credentials can be used, so be careful how they're shared before they're used. Your org API key is a different matter entirely. It's admin-grade. Keep it on your server, keep it out of logs, and rotate it through the console if you think it's been exposed. The create route is never CORS-enabled and we never accept an org key from a browser, so the only way it reaches one is if your own code puts it there. ## Tracking usage with attribution The attribution tags you send aren't just labels. Those values surface again through our observability, so you can see what users of your embedded agents are doing, broken down by whatever dimensions matter to you. ```json "attribution": { "plan": "enterprise", "region": "emea", "team": "support" } ``` We don't inspect what you put in them. Up to 16 keys, keys up to 64 characters, values up to 256 characters. These limits are enforced, and a request that breaks them is rejected rather than quietly truncated, so you never end up with a session carrying tags you think are there but aren't. ## Errors to handle | What happened | Status | Code | |---|---|---| | A field in the body that isn't in the table above | 400 | `validation_error` | | Both email and externalId, or neither | 400 | `validation_error` | | An email value with no `@` in it | 400 | `validation_error` | | An externalId that breaks the format or length rule | 400 | `validation_error` | | An unknown user, without `createUserIfNeeded` | 404 | `not_found` | | The agent isn't one of your organization's agents | 404 | `not_found` | | Your key's organization doesn't match the orgSlug | 404 | `not_found` | | The envSlug isn't your key's Environment, or doesn't exist | 404 | `not_found` | | Your key doesn't carry the scopes this route needs | 404 | `not_found` | The 404s are deliberate and they're all identical. A wrong organization, a wrong Environment, a wrong agent and an insufficient key all answer the same way, so nobody holding a key can map out which organizations or Environments exist by probing. It does mean a 404 during setup needs checking against all four, since the response won't narrow it down for you. ## Things that catch people out Query parameters are ignored, extra body fields are rejected. Addressing is entirely in the path. The asymmetry is deliberate but surprising: a stray query string does nothing, a stray body field is a 400. The embed always runs the agent's current published config. Publish a new version and your users pick it up on their next turn. A turn already in flight finishes on the version it started with. The credential is end-user tier no matter who the user is. There's no admin session credential. `createUserIfNeeded` defaults to false. Most first integrations want it true, and the symptom of leaving it out is a 404 that looks like the agent is wrong. --- ## The element **Section:** SDK reference / undefined The drop-in half of the SDK. One custom element that renders a complete chat interface on your page, carrying its own styles, in its own shadow root. Your page loads one script and writes one tag. ## Loading it ```html ``` The script registers the `mindset-agent` tag itself. ## Configuring it `configure()` is what makes the element live. Until you call it, the element renders nothing and `send()` won't work. ```javascript element.configure({ getSession: () => myBackend.fetchAgentSession(), }); ``` `getSession` is a function that returns the session credential, not the credential itself. Your backend creates that credential server-to-server; the element decodes it, runs on it, and renews it without your involvement. Passing a static string won't work, because a string can't be renewed. An organization API key is dropped rather than sent. The credential carries everything else the element needs, including your organization and Environment, so there's nothing further to configure. ## Attributes and properties Every attribute mirrors a property in both directions, so declarative and imperative configuration can't disagree. React 19 sets unknown props as DOM properties, which is why each one is a real accessor pair. | Property | Attribute | What it is | |---|---|---| | `agent` | `agent` | Which agent to run. Its handle, or its UUID. Copy either from the agent's Embed tab | | `conversationId` | `conversation-id` | Which conversation to resume. See below | Note the attribute name is `conversation-id` in markup and `conversationId` as a property. HTML lowercases attribute names when it parses, so the two forms are unavoidable. ## Resuming a conversation By default, every time the element mounts your user gets a fresh conversation. If you want someone to come back to what they were doing, this is the section that matters. It's a round trip, and you don't need to know an ID to start using one. **First visit.** Leave `conversation-id` off. The element starts a fresh conversation. On the first completed turn, once there's something worth coming back to, it tells you the ID two ways: it fires a `mindset:conversation` event, and it writes the ID onto its own `conversation-id` attribute. Store it against whoever is looking, in your own database, keyed to that user. **Next visit.** Set `conversation-id` to the value you stored. The agent picks up with the history it had. ```javascript element.addEventListener("mindset:conversation", (e) => { saveForCurrentUser(e.detail.conversationId); }); ``` ### Rules worth knowing Prefer the event over reading the attribute back. If your stack morphs the DOM towards a template, it will strip an attribute the template doesn't declare, and the element deliberately doesn't fight your renderer to put it back. Store what the event gave you. Clearing the attribute does nothing. That's on purpose. An unrelated DOM update must not cost a user their conversation. Writing the same ID back is safe. A framework re-render doing exactly that is a no-op. Writing a different ID is a reconfiguration. The running conversation stops, loudly, rather than switching silently. ### One thing to get right A conversation ID is per user. Resolve it per request, from what your backend knows about who is asking. Templating one into static server-rendered HTML serves one user's conversation to everybody who loads the page, and no check inside the element can tell that apart from a host legitimately setting the attribute from code. The ID is a correlation key, not a credential. It already travels in every request the element makes, and holding it grants nothing without the session credential. That's why it's safe to express as an attribute at all. ## Methods | Method | What it does | |---|---| | `configure()` | Set this instance up. Calling it again tears down the previous conversation and starts one for the new config, so a host that reconfigures on identity change can't end up with two live conversations on one element | | `send()` | Send a user turn. Before `configure()`, this reports an error rather than silently dropping the message | | `stop()` | Stop the turn in flight. Safe to call when nothing is running | Each of these maps onto a documented headless command. The element is a thin shell over the same contract, with no private channel to the runtime. ## Events DOM CustomEvents, namespaced so they can't collide with your own, dispatched so a listener on an ancestor outside the shadow root still sees them. | Event | detail | |---|---| | `mindset:runtime-event` | A RuntimeEvent. The full runtime vocabulary, surfaced directly. See the events reference | | `mindset:error` | Carries `code`, `message` and an optional `cause`. Configuration and transport failures you should surface | | `mindset:conversation` | Carries the `conversationId`. Fires once: at mount if you supplied the ID, or on the first completed turn if the element started fresh | `mindset:conversation` carries the ID and nothing else, deliberately. It crosses the shadow boundary into your document, where every script on the page can read it. These are additive only. A new event type is a compatible change; renaming or reshaping one is not. Write consumers that switch on the type and ignore anything they don't recognize. ## Error codes `mindset:error` detail is always structured, never prose alone. Branch on `code` and treat an unknown code as a generic failure. | code | Raised when | |---|---| | `missing-agent` | `configure()` ran with no agent, from either the attribute or the config | | `invalid-auth` | `configure()` got no `getSession` function | | `not-configured` | `send()` was called with no live conversation, either before `configure()` or after a real disconnect | | `stale-configuration` | An attribute changed after `configure()`, so the live conversation no longer matches what the element reports. Call `configure()` again | | `turn-failed` | The turn itself failed: transport, credentials, or the run | `turn-failed` is worth handling specifically. It's the class of failure you can't diagnose any other way, because the SDK fetches its envelope before the first runtime event, so an expired credential produces no `mindset:runtime-event` at all. ## Theming Theming is CSS custom properties that pierce the shadow boundary. There's no `::part()` API in this version, so don't plan around one. The element ships a default set of colors on its own shadow host, so an embed on a page with no theme renders in real colors with no cooperation from you. Add `class="dark"` to the tag for the dark set. Override any channel from your own stylesheet. A rule on the host element wins over the element's internal defaults. ```css mindset-agent { --ch-accent: 124 58 237; /* space-separated r g b */ --ch-page: 255 255 255; } ``` Values are space-separated RGB components, not hex and not `rgb()`. The full vocabulary is 62 channels. The list may change as the interface develops, so treat it as the current set rather than a fixed contract. ``` --ch-page --ch-panel --ch-card --ch-surface --ch-input --ch-card-hover --ch-selected --ch-overlay --ch-text-heading --ch-text-primary --ch-text-secondary --ch-text-muted --ch-logo --ch-edge --ch-edge-subtle --ch-edge-strong --ch-btn-primary --ch-btn-primary-hover --ch-btn-secondary --ch-btn-secondary-hover --ch-accent --ch-accent-hover --ch-on-accent --ch-agent --ch-agent-hover --ch-node-agent --ch-node-connection --ch-node-tool --ch-node-knowledge --ch-node-widget --ch-node-function --ch-chip --ch-chip-text --ch-chip-border --ch-status-warning-bg --ch-status-warning-text --ch-status-warning-border --ch-status-success-bg --ch-status-success-text --ch-status-success-border --ch-status-info-bg --ch-status-info-text --ch-status-info-border --ch-status-danger-bg --ch-status-danger-text --ch-status-danger-border --ch-status-neutral-bg --ch-status-neutral-text --ch-status-neutral-border --ch-phase-specify-bg --ch-phase-specify-text --ch-phase-specify-border --ch-phase-specify-container --ch-phase-build-bg --ch-phase-build-text --ch-phase-build-border --ch-phase-build-container --ch-phase-verify-bg --ch-phase-verify-text --ch-phase-verify-border --ch-phase-verify-container --ch-phase-draft-container ``` Isolation runs both ways and it's structural, not conventional. The interface lives in a shadow root and the compiled stylesheet is adopted into that root rather than added to your document, so neither side's CSS can reach the other. --- ## The UI-less SDK **Section:** SDK reference / undefined The same agent, the same conversation, with no interface at all. You get typed events as the agent works and commands to drive it, and you render every pixel yourself. It is the same underlying code for our agent, but designed so you can build your own components around it. Use it when the conversation has to live inside an interface you already have, when the agent's output drives something other than a chat transcript, or when your design system won't tolerate a component you don't control. If our chat interface is fine, use the drop-in element instead. It's a thin shell over exactly this contract. ## Loading it The SDK is an ES module served from your Mindset AI host. There's no package to install. ```javascript const { createAgentConversation } = await import( "https://YOUR-MINDSET-HOST/sdk/mindset-agent-uiless.js" ); ``` ## Creating a conversation ```javascript const chat = createAgentConversation({ agent: "support-bot", getSession: async () => { const r = await fetch("/api/session"); // from your application back end if (!r.ok) throw new Error("Could not start a session"); const { session } = await r.json(); return session; }, }); ``` | Option | Type | What it does | |---|---|---| | `agent` | string | Required. Which agent this conversation talks to, by handle | | `getSession` | function | The credential provider. Returns the opaque session your backend created | | `conversationId` | string | Resume an existing conversation. Leave it out for a fresh one | `getSession` is a function, not a string. The SDK calls it, decodes the credential, presents it on every call, and renews the session itself. When the session's outer lifetime is spent it calls your function again for a fresh one. You write no refresh logic. A static string can't be renewed, so it isn't accepted. An organization API key is dropped rather than sent, and the call will fail with a 401 rather than leaking your key into a browser-reachable header. Because the credential carries the organization, the Environment and the platform origin, `org` and `baseUrl` are both unnecessary on a normal integration. ## Driving the conversation | Command | What it does | |---|---| | `send(text)` | A user-typed turn. Resolves to the agent's reply | | `sendMessage(text, options)` | A turn your application triggered rather than the user. Pass `{ silent: true }` to keep the trigger out of the displayed transcript while the agent still receives it | | `widgetAction(action)` | Send a rendered widget's action payload back into the conversation. Framed as text and driven as a silent turn, so the agent reasons about the interaction without a stray user message appearing | | `stop()` | Cancel the turn in flight. The model call is cut and the spend stops | | `retry()` | Re-run the last turn, replacing its reply. Does nothing if there isn't one | | `reset()` | Clear all history and start fresh on the same agent | `send` and `sendMessage` both resolve to the reply as a string, so you can await them if that suits your interface. Most interfaces render from events instead, since awaiting means waiting for the whole turn. `stop()` ends the turn with a `run_error` carrying `aborted: true`. That's a cancellation, not a failure. Render it as "stopped". ## Listening ```javascript const unsubscribe = chat.on((event) => { switch (event.type) { case "text_delta": append(event.content); break; case "complete": finish(event.response); break; default: break; // ignore anything you don't recognize } }); ``` `on()` returns a function that unsubscribes. Two families of event arrive here: the runtime's vocabulary, and the SDK's own events about the conversation. Both are covered in the events reference. Keep a default arm in your switch, because the vocabulary is additive and new types will appear. ## Reading state | | What it gives you | |---|---| | `messages()` | What the model sees. Includes silent triggers, the assistant turns that called tools, and the tool results they produced | | `transcript()` | What you should display. User and assistant turns, without silent triggers and without tool rows, plus any rendered widgets on `.widget` | | `conversationId` | The ID this conversation's turns are grouped under. Read it to correlate with your own systems | | `envelope()` | The session envelope, fetched once and cached for this instance | Two of these need care. `transcript()` is not free. It re-parses every widget envelope in the conversation on each call. Fine to call when state changes, but hold the result rather than calling it on every render. If you're painting restored history, the `history_settled` event hands you the same list already built. `messages()` and `transcript()` are two different lists, not two views of one. If you render from `messages()` your users will see silent triggers and raw tool traffic they were never meant to. ## Talking to the agent from your page Three commands connect your application's state and capabilities to the agent. | Command | What it does | |---|---| | `setPageTools(tools)` | Functions in your page the agent may call. Replaces the whole set | | `setSituationalAwareness(entries)` | Facts about what the user is currently looking at, folded into the agent's context each turn | These are covered properly in "Host and agent channels". One thing belongs here, because it's the single most consequential fact about this SDK. ### Page tool handlers run under adversarial influence The model decides whether to call your tool and what arguments to pass it. The model is influenceable by anything in its context, which includes your knowledge base and the output of other tools. Two consequences people get wrong. The `jsonSchema` you declare is not an input filter. It's guidance to the model about what to send. Nothing validates the arguments against it before your handler runs. Validate them yourself. There is no server-side re-authorization. An in-page function can't be re-authorized by us, which is exactly what makes it a page tool. Nothing checks that this user is allowed to do this thing at the moment the handler runs. So treat every argument as attacker-controlled, and put nothing behind a page tool that you wouldn't expose to a hostile caller. No mutations, no payments, no reading personal data, unless your handler does its own authorization first. What limits the damage is that you choose what to expose. Your organization API key never reaches the browser, and privileged operations still re-authorize on our servers. --- ## Events reference **Section:** SDK reference / undefined Everything the agent tells you while a turn runs, and after it finishes. ## Two vocabularies, one channel Two different sets of events arrive on the same listener, and they come from different places. Runtime events are what the agent runtime emits as it works: text arriving, tools running, the turn finishing. They're the bulk of this page. SDK events are about the conversation itself rather than the turn: which conversation this is, and what history was restored. The runtime knows nothing about stored conversations, so the SDK owns these. You handle both in the same listener. The distinction matters when you're looking something up, because they're documented separately, and it matters if you're using the drop-in element, because only one of the two crosses into your DOM. ## How you receive them With the drop-in element, runtime events arrive as `mindset:runtime-event` DOM events, with the runtime event as `detail`. ```javascript element.addEventListener("mindset:runtime-event", (e) => { const event = e.detail; if (event.type === "text_delta") append(event.content); }); ``` With the headless client, both vocabularies arrive on `on()`. ```javascript chat.on((event) => { if (event.type === "text_delta") append(event.content); }); ``` Configuration and transport failures are a separate channel. On the element they're `mindset:error`, and they carry their own codes rather than being runtime events. ## The rule that keeps your integration working Switch on `type` and ignore anything you don't recognize. This vocabulary is additive only. New event types will appear, and existing ones won't be renamed or reshaped. A consumer that handles what it knows and skips the rest keeps working across updates. A consumer that throws on an unfamiliar type will break on one. ## The shape of a turn `run_started` is the first event of a run. `run_finished` or `run_error` ends it. `complete` carries the final answer. What happens in between varies by turn and by model. Don't build logic that assumes a fixed order beyond those brackets. ## Runtime events | `type` | `detail` | What it means | |---|---|---| | `run_started` | `{ runId? }` | The turn has started | | `run_finished` | `{ runId? }` | The turn finished cleanly | | `run_error` | `{ message, runId?, aborted?, reason?, effects?, stepLimit? }` | The turn ended in an error. Terminal. See abort semantics below | | `text_delta` | `{ content }` | A chunk of the reply. Concatenate `content` across all of these for the full text. Each one may be a fragment | | `thinking_delta` | `{ content }` | A fragment of the model's summarized reasoning, while it thinks | | `stream_flush` | `{}` | Token streaming ended for the current chunk. Buffering displays flush here. Not the same as `complete`, which means the turn is done | | `tool_announce` | `{ toolName, toolCallId }` | The model has named a tool it's about to call, while its arguments are still streaming | | `tool_start` | `{ toolName, toolCallId, args? }` | A tool is about to execute. `toolCallId` pairs with the matching `tool_end` | | `tool_end` | `{ toolName, toolCallId, durationMs, output?, isError? }` | A tool finished. `isError` marks a failed call, with the error in `output` | | `widget-payload` | `{ doc, kind?, summary?, toolCallId? }` | A renderable payload produced during the turn. Read the note below before using it | | `citation_applied` | `{ enrichedText }` | The reply with citation markers inserted | | `quick_replies` | `{ question, options }` | Two to four tappable canned responses under a short prompt | | `follow_up_questions` | `{ questions }` | Two or three open-ended prompts to continue with | | `references` | `{ records }` | The turn used retrieval and produced source references | | `conversation_title` | `{ title }` | A generated short title for the conversation | | `complete` | `{ response, messages, threadId?, bound? }` | The final answer. `response` is the agent's text | ### Three that need care `thinking_delta` is display only. It never becomes part of the answer and it isn't stored with the conversation. Models that don't think never emit it, so tolerate its complete absence rather than waiting for it. `tool_announce` is a timing hint, not a guarantee. It only fires when the model adapter can see the provider's stream boundaries. `tool_start` always fires when execution actually begins, correlated by the same `toolCallId`. Build on `tool_start`. `widget-payload` needs you to branch on `kind`. Don't hand `doc` straight to a renderer. For the presentational kinds (`chart`, `callout`, `table`, `card`, `badge`, `stat` and `widget`), `doc` is a widget tree you render through our widget renderer. For the interactive kinds (`quick_replies` and `capability_cards`), `doc` is a different shape entirely, carrying `replies` or `cards`. These are meant to be drawn as tappable things that send the chosen option. Passing them to the renderer loses the interaction. Like `thinking_delta`, this one is display only and won't appear on every turn. ## Abort semantics When someone stops a turn, you get `run_error` with `aborted: true`. That marks a cancelled run rather than a failure: the caller's abort signal fired, the model call was cut, and the loop exited. It's a requested outcome. Render it as "stopped", not as an error. It's a marker on `run_error` rather than its own event type so that consumers switching exhaustively over the vocabulary don't have to change. ## SDK events These arrive on `on()` alongside the runtime vocabulary when you're using the headless client. ### `conversation_id` Which conversation this is. Fires once, either at mount when you supplied the ID or on the first completed turn when the SDK started a fresh one. Store it against whoever is looking and hand it back on their next visit. That round trip is the whole capability. It's a correlation key, not a credential. It grants nothing without the session credential. The drop-in element re-emits this one to your DOM as `mindset:conversation`. ### `history_settled` The restored transcript. This is how a custom interface paints a returning user's history. Fires once per conversation, after hydration settles, and it always arrives. Restored, empty, or failed, you get the event, so you can safely hold your composer closed until it lands. The read carries a fifteen-second bound, and a connection that stalls is dropped to a cold start rather than hanging forever. That bound is a liveness guarantee, not a display policy. It's generous on purpose and shouldn't fire on a working connection. Set your own shorter hold for how long you'll keep the interface waiting. The payload is display-safe as it arrives, and frozen, so one listener can't corrupt what another sees. It's the prose turns plus any rendered widgets those turns displayed, with tool traffic filtered out. Each message carries an optional `widget` of `{ doc, summary, kind }`. You can ignore it. A rendered turn's `content` is that widget's own plain-text summary, so reading `{ role, content }` alone still paints an honest transcript. Read `widget` only if you want to redraw the original. Two things it can't tell you. It doesn't distinguish "this user has no history" from "we couldn't load it", because a failed read is a silent cold start. So you can't render "couldn't load your history" from this event. And it's deliberately not a DOM event. The drop-in element consumes it internally and doesn't rebroadcast it, because DOM events cross the shadow boundary into your document where any script could read them, and this is the user's private conversation content. If you're using the element you won't see this event, and you don't need to: the element has already painted the history. Don't build display for these yet. Do tolerate them, under the same rule as any unrecognized type. Note that `interrupt` is not used for cancellation, despite the name. A cancelled run is `run_error` with `aborted: true`, as above. --- ## Use your own model provider keys: AWS **Section:** Deploy / undefined Your organization can run its agents on your own model provider account rather than ours. You supply the credential, the provider bills you directly, and with AWS Bedrock the inference itself runs inside your own AWS account, in an AWS region you pick. An org admin sets this up once, in Settings → Model keys. It applies to every agent in your organization. ## What this covers Setting a model key changes two things: whose account pays for inference, and which endpoint serves it. For AWS Bedrock it also changes where the inference physically runs. It doesn't change which models you can pick, how your agents are built, or where the rest of your data lives. Your organization's data stays on the Mindset AI deployment you're on. ## Supported providers | Provider | What you supply | |---|---| | Anthropic | An API key | | OpenAI | An API key | | Google | An API key | | Open-weights hosts | An API key for any OpenAI-compatible or self-hosted endpoint | | AWS Bedrock | A cross-account IAM role, plus the AWS region to run in | Bedrock is the one that works differently, because it doesn't use a bearer key at all. It has its own section below. ## Before you start You'll need to be an org admin. Model keys are configured in the Mindset AI console only. There's no API, no SDK method and no agent tool for setting or reading them, which is deliberate: it removes a whole class of credential exposure rather than guarding against it. You'll also want the provider account ready first, since the console asks you to paste the credential in one go. ## Set a provider key For Anthropic, OpenAI, Google and open-weights hosts, the flow is the same. 1. Go to Settings → Model keys. You'll see a row for every supported provider, whether or not you've configured it. 2. On the provider you want, choose Set key. Each row links straight to where that provider's console issues keys. 3. Paste the credential and give it a label, something like "Production Anthropic". The label is how you tell your keys apart later. 4. Save. We run a quick check against the provider when you save. If it succeeds, the row shows as verified. If the provider is slow or unreachable, the key is still saved and the row shows as unverified with a note about what the check failed on. An unverified key still works at runtime, so a provider outage never blocks you from saving. You'll never see the credential again after you save it. More on that below. ## Set up AWS Bedrock Bedrock runs Claude models on your own AWS account. Rather than handing us a long-lived AWS access key, you create a role in your account that trusts ours, and we assume it for the length of each session. ### Create the role in your AWS account In the AWS account you want the inference billed to, create an IAM role with: - A trust policy naming the Mindset AI principal, with an external ID you choose. - A permissions policy scoped to `bedrock:InvokeModel`. The external ID is a shared secret between your trust policy and your Mindset AI configuration. It's what stops anyone else persuading us to assume your role on their behalf, so pick something unguessable and treat it as a credential. Your trust policy will look roughly like this: ```json { "Version": "2012-10-17", "Statement": [ { "Effect": "Allow", "Principal": { "AWS": "" }, "Action": "sts:AssumeRole", "Condition": { "StringEquals": { "sts:ExternalId": "your-chosen-external-id" } } } ] } ``` Make sure the models you want are enabled on Bedrock in the region you're going to pick. Bedrock model access is granted per region in your own AWS console. ### Add it in Mindset AI 1. Go to Settings → Model keys and choose Set key on the Bedrock row. 2. Enter the role ARN from the step above. 3. Enter the external ID your trust policy requires. 4. Enter the AWS region you want inference to run in, for example `eu-west-1` or `us-east-1`. There's no default and nothing is pre-selected, so you have to state it. 5. Confirm the acknowledgement about where your inference will run. 6. Save. The region field is free text rather than a dropdown. That's on purpose. A curated list would imply we'd vetted those regions for your compliance position, and we haven't. You know your requirements and we don't, so the choice is yours to state. An unrecognized region is refused rather than quietly replaced with a default. ### What we do with the role Each time your agents need to run, we make a single `sts:AssumeRole` call against your role and get back temporary credentials that last one hour. We cache them until shortly before they expire, then get fresh ones. Nothing long-lived is ever stored. In your own CloudTrail, these appear under the session name `mindset-bedrock`, so you can audit exactly what we did and when. We only ever assume the role you named. The permissions your role grants are the outer limit of what we can do in your account. ## Where your data goes With Bedrock configured, your organization has two answers to "where does our data go", and they're both true: - Your inference runs on your AWS account, in the AWS region you named. - Everything else stays on the Mindset AI deployment your organization lives on. Those can be different places, and that's the point of the feature. A few things to be clear about, because they're the questions a compliance review will ask. The region is your choice, not a boundary we police. We don't keep a list of which AWS regions are acceptable for which customers, and we don't check your choice against one. We're not in a position to know your legal requirements, and a list that looked authoritative would be worse than no list at all. What we do instead: you pick the region explicitly, you acknowledge it, we record who made that choice and when, and the region stays visible on the panel afterwards so whoever inherits your organization can see it without re-running the setup. Retrieval runs on a separate key. Searching your knowledge for relevant context is a different call from generating a reply, and it uses a different credential. You can set an embedding key for your own organization under Settings → Embedding keys. That key is OpenAI only, and the embedding model is fixed at `text-embedding-3-large`. Neither is arbitrary. Anthropic has no embeddings API at all, and Google and Bedrock embed into a different vector space, not a wider or narrower one, so results from a different provider wouldn't be comparable to what's already stored. Your whole knowledge base is indexed into one space, and every search has to land in that same space to mean anything. So configuring Bedrock covers your chat and completion traffic. It doesn't move your retrieval traffic to AWS, and no setting available today would. One key per organization, across every Environment. A model key belongs to your organization, not to a particular Environment within it. If you keep separate Environments for testing and production, they share the same key. There's no way today to run a sandbox key in one and a production key in another. ## Replacing and removing a key We can't show you a key back. Not in the console, not through an API, not in a log line, not in an error message. Once you save it, the value has left our servers and only a reference remains. This applies to every provider, and to your Bedrock external ID. That means replacing a key means typing the whole thing again. The form is never pre-filled, because there's nothing to pre-fill it with. It's friction, and it's the same property that makes the credential safe. Removing a key takes effect immediately. For Bedrock, we also drop any cached session, so a removal is a real revocation rather than one that takes effect an hour later. Be careful with this one. If you remove your only configured key, your organization stops being able to run agents at that moment. Admins will see a message pointing them back to Settings → Model keys. Anyone else will see a note saying the organization isn't set up yet, rather than a chat box that accepts a message and then fails. ## Where this is available Model keys, including Bedrock, are available on our integration and EU deployments today. US production is planned, and model keys will be available there when it launches. | Mindset AI deployment | Model keys available | |---|---| | Integration | Yes | | EU | Yes | | US | Planned | This is about where your Mindset AI organization is hosted, which is separate from the AWS region you choose for Bedrock inference. You can run Bedrock inference in any AWS region from either of the deployments above. ## Troubleshooting **The Bedrock row saved but shows as unverified.** Our check couldn't assume your role. The note on the row names what failed. The usual causes are a trust policy that doesn't name the Mindset AI principal, an external ID that doesn't match what you entered, or a role without `bedrock:InvokeModel`. **Agents fail with an assume-role error after working fine.** Something changed on the AWS side. Check the role still exists, its trust policy is intact, and the external ID hasn't been rotated in AWS without being updated in Mindset AI. **A model isn't available.** Bedrock model access is granted per region in your AWS account. Check that the model is enabled in the region you configured. **Error messages never contain your credentials.** If you're reporting a problem to us, our errors name the field that failed and the provider's own refusal, never the value. You can share them safely. **An admin needs to change this and can't find it.** Model keys are org-admin only and live under Settings → Model keys. Members can't see the panel. --- ## Mindset AI in your own AWS account **Section:** Deploy / undefined The Mindset platform runs as a self-hosted deployment inside your AWS environment, on your model keys, behind your network controls. Your data, your agents and your configuration stay where you put them. This page describes what that deployment looks like, what we need from you to build it, and how it stays current afterwards. This describes the architecture and set up process for self-hosted AWS deployments. If you are planning a deployment, talk to us and we will confirm the specifics for your environment. ## 01. The stack ### What it is built on One container image, a managed Postgres database and a handful of standard AWS services. Nothing exotic, nothing that needs a platform team to learn a new toolchain. The whole platform ships as a single container image: the API, the agent runtime and the web application in one process. That is a deliberate constraint rather than a simplification: one artefact to deploy, one thing to roll back, and no orchestration between internal services for your team to reason about or monitor. It runs on ECS Fargate, so there are no servers or clusters in your account to patch or scale. Tasks come and go against your load, AWS maintains the underlying host, and the deployment is defined entirely by a task definition you can read. We deliver the infrastructure as a CloudFormation template. You own it and you run it. There is no state file to host, no extra binary to install and no tool version to keep in step with ours. You deploy a template from the console or the CLI, and AWS tracks the stack, reports drift and rolls back a failed change on its own. ### Agents are the interface, not a feature Every capability in the platform is callable by an agent as well as by a person, because both go through the same service layer. The same surface is exposed over MCP, which means your own tooling and your own agents can drive the platform directly, and it is how your data, configuration and the agents you build are exported whenever you want them, without asking us. ## 02. The environment ### What the deployed environment looks like A private VPC in your account and your region. Nothing is publicly reachable except the load balancer, and traffic only ever leaves towards the model providers you have chosen. | | | |---|---| | **Application** | TypeScript on Node. One image, one process, serving both the API and the web application. | | **Data** | PostgreSQL on Amazon RDS, with vector search for retrieval. A second, separate database holds the append-only audit record. | | **Objects** | Amazon S3 for conversation history and generated assets, encrypted with your own KMS keys. | | **Credentials** | AWS Secrets Manager. Model keys and integration credentials are held there and referenced, never stored in the database. | | **Edge** | An Application Load Balancer with AWS WAF in front, an ACM certificate on your own hostname, and access logs delivered to your S3. | | **Models** | Your provider accounts and your keys: Amazon Bedrock, Anthropic or OpenAI for inference, and your choice of embedding provider for retrieval. | | **Sign-in** | Google sign-in or email and password. Federation with your own identity provider is on our roadmap. | Every request enters through WAF and the load balancer, which is the only source the application tasks accept. The tasks sit in private subnets with no public address, and the only traffic leaving your account goes to the model providers you configured. ![Architecture diagram: users reach the VPC over HTTPS through AWS WAF and a public-subnet load balancer, which is the only allowed source into the private app subnet's Fargate tasks; one task reads and writes the application RDS Postgres store, the other appends to the audit RDS Postgres store, writes objects to S3, reads credentials from Secrets Manager, and reaches a NAT gateway for outbound-only egress to model providers (Bedrock, Anthropic, OpenAI); every store is encrypted with your own KMS key; release images arrive by a one-way push from the Mindset release channel into your own ECR repository, with no outbound connection back to Mindset.](/images/docs/aws-deployment-architecture.png) The connections above describe the kinds of connection rather than a partition: either task reaches any store. Release images are pushed inbound to your own registry; nothing calls back out to us. ### Architecture: what lands in your account | Resource | Purpose | Notes | |---|---|---| | VPC + subnets | Private network for the application and databases | Public subnets carry only the load balancer and NAT | | ECS Fargate service | Runs the platform container | No servers or clusters to manage | | Application Load Balancer | Terminates TLS, routes to the tasks | The only publicly reachable component | | AWS WAF web ACL | Managed rule groups and rate limiting at the edge | Attached to the load balancer; you tune the rules | | ACM certificate | TLS on your own hostname | DNS-validated in your zone | | RDS PostgreSQL | Application data and vector retrieval | Not publicly accessible | | RDS PostgreSQL (audit) | Append-only audit record | Separate store, separate database role | | S3 bucket | Conversation history and generated assets | Public access blocked, versioned | | S3 bucket (logs) | Load balancer access logs | Retention is yours to set | | Secrets Manager | Model keys and integration credentials | Access scoped by resource prefix | | KMS key | Encryption at rest across every store | Your key, your key policy, your revocation | | ECR repository | Holds the platform images you deploy | We push, you deploy | | CloudWatch log group | Application logs | Never leaves your account | | IAM roles | Task execution and task runtime identity | Least privilege, scoped to the above | ### The security posture, by default The template ships with the controls already in place, rather than as recommendations in a runbook. Application tasks have no public address and accept traffic only from the load balancer's security group. WAF managed rules and rate limiting are attached at the edge, with access logging on so that what the edge did is a matter of record. Every store is encrypted with a key you control, and the audit record lives in a separate database with a role that can append but cannot rewrite. Because the whole environment is a CloudFormation stack, what protects it is visible in your own tooling rather than something you take on trust. The stack records exactly what was deployed, and AWS drift detection reports anything changed outside it, so your team can audit the security position directly, on your own schedule, without asking us. ### What does not happen - **No call-home.** The deployment makes no connection to Mindset infrastructure. Nothing about your usage, your content or your configuration is reported to us. - **No licence kill switch.** Images live in your registry and run from your account. Your ability to run the platform does not depend on reaching us. - **No Mindset control plane in the request path.** No component of ours sits between your users and your deployment. - **No data leaves for us.** Your data, configuration, content and the agents you build stay in your environment, and all of it is exportable at any time through the API and MCP. ### Regions and residency A deployment lives in one AWS region and holds its own data. Where you need data to stay in more than one jurisdiction, you run more than one deployment: each independent, each in its own region, with no shared database between them. AWS charges for all of it go directly to your account, on your own commercial terms with AWS. ## 03. Getting it built ### How we get it built with you We work alongside your team rather than handing over a document. You provision in your own account, with us in the room. Standing the deployment up is the first stage; building your first use case on top of it is the next. ### What we need from you Five things, and they are the whole list. Everything else is ours. | | | |---|---| | **An AWS account you control** | In the region you want the deployment to live in, with someone able to provision networking, databases and IAM. A dedicated account is cleanest, but not required. | | **A container registry and a role to push into it** | An ECR repository in your account, and one cross-account role that lets us push signed builds to it. That is how releases reach you, and it means nothing at runtime depends on us. | | **A hostname and the ability to change DNS** | The platform runs on your own hostname. We need it decided early, because sign-in redirect addresses are matched exactly and changing the hostname later means reissuing them. | | **Your model provider credentials** | Keys for your chosen inference provider and your chosen embedding provider. You hold them, you set the spend limits, and the providers bill you directly. | | **A named engineer for the setup window** | One person with provisioning rights, reachable in a shared channel while we stand it up. This is consistently what decides whether a deployment takes days or weeks. | The split matters more than the sequence: we supply the template, the images and the expertise, and your team performs every action inside your own account. Nobody from Mindset needs standing access to your environment for a deployment to happen. Note where this flow ends: a live, validated deployment. Building your first use case starts from there and is its own stage of work. ### The stages ![Stage diagram: four stages, Prepare, Provision, First deploy, Validate, each with a Mindset row and a Your team row running in parallel, converging on "Deployment live and validated" and then the next stage, "Your first use case". Caption: your team holds every credential and performs every action inside your account throughout.](/images/docs/aws-deployment-stages.png) | Stage | Mindset | Your team | |---|---|---| | Prepare | Template and requirements | Account, registry, DNS, model keys | | Provision | Paired with your team, live | Deploy the template | | First deploy | First image into your registry | Run the deploy, apply migrations | | Validate | Smoke checks with your team | Confirm and sign off | | *Deployment live and validated* | → | *Next stage: your first use case* | Your team holds every credential and performs every action inside your account throughout. 1. **Prepare.** We confirm the shape for your environment, send the template and the exact IAM policies we need, and agree what the first use case is and how you will judge it. 2. **Provision.** Your team creates the account prerequisites and deploys the template, with us alongside in a shared channel. Nothing about this step requires access for us. 3. **First deploy.** We push the first signed image to your registry. You run the deploy and the migration step, both inside your own network. 4. **Validate.** Smoke checks against the running deployment alongside your team, so it is confirmed working in your environment rather than assumed from ours. 5. **Then: your first use case.** With the deployment live, the work moves inside the product: configuring and testing your first use case, with our enablement team alongside yours. That is its own stage of work, and it is where the value actually shows up. ## 04. Staying current ### How it stays updated and supported We ship releases to your registry. You decide when to apply them. Nothing upgrades itself, and nothing requires us to reach into your account. A release is not just a new image, so we do not ship one. Each release arrives as a bundle: the image, the infrastructure template version it expects, the database migrations that go with it, notes on anything an operator has to do, and a clear statement of whether it can be rolled back. Those parts are versioned together, so there is never a question of which template goes with which image. ### The release path ![Release path diagram: Mindset build (tested, versioned, signed: one build, the same version for every deployment) produces a release bundle (image version, template version, migrations, notes), pushed into your AWS account's ECR, current plus previous versions held. You decide when to deploy; on deploy, migrations run inside your VPC; the new version starts serving. A dashed arrow shows rolling back by redeploying the previous version, which never left your registry.](/images/docs/aws-deployment-release-path.png) | Step | What happens | |---|---| | Mindset build | Tested, versioned, signed. One build, the same version for every deployment, so they stay comparable | | Release bundle | Image version · template version · migrations · notes | | Push → your ECR | Current + previous versions held in your own registry | | You decide when to deploy | Every decision after the push is yours | | On deploy → migrations | Run inside your VPC | | New version now serving | To roll back, redeploy the previous version: it never left your registry | The arrow crosses the boundary in one direction only: we publish into your registry, and every decision after that is yours. Rolling back is not a special procedure: the previous version is still in your own registry, so you redeploy it. ### Upgrading on your own schedule You deploy the release you want, when you want it. Database changes are designed so that the version you are running keeps working against the migrated schema, which is what makes it safe to roll the image back without touching the database. Where a release cannot be rolled back, the notes say so before you start rather than after. We support a defined window of recent releases and test the upgrade path across it, so skipping a release is expected rather than risky. If you have stayed further back than that window, we will get you current. We would just rather tell you that plainly than let you discover it during an incident. ### Support Maintenance, patching, version upgrades and bug fixes are part of the subscription, and self-hosting does not change that. What we warrant is the software and its maintenance: the infrastructure it runs on is yours, and its uptime is in your hands, which is the honest consequence of it being your account. In practice, support happens in a shared channel with direct access to the people who build the platform, alongside regular reviews covering usage, roadmap and what to build next. Because the deployment does not report to us, we will ask you for logs when we are diagnosing something, a deliberate trade for the fact that nothing about your environment is visible to us by default. --- --- # Getting Started ## What Mindset is **Section:** Start here ## The one idea **We turn the skills, MCPs and ideas your team has already built into reliable agents your company owns.** Right now that work is on laptops. Someone in finance wrote a skill that reconciles invoices. Someone in ops has an MCP server wired to the ticketing system. Someone else has a prompt they run every Monday that everybody now depends on. Each one helps. Together they are a problem, because: * It is not centralised. It lives with whoever built it. * Only one person can fix it, or you pull an engineer off real work. * You cannot prove what it did, or what it cost. Mindset takes what you have already got and runs it somewhere the company owns: one place to build agents, one place to see what they did, one set of connections they are allowed to touch, and no dependency on a single model vendor. ## The example we use throughout Every article in this guide uses the same workflow, so you see one thing from every angle. **Invoice exceptions.** Every month, finance gets invoices that do not match the purchase order. Someone works through the list by hand: pull the invoice, find the PO, work out what the difference is, check whether the contract allows it, and either let it through or query it with the supplier. By the end of this guide, an agent does the first four steps and a person approves the fifth. ## The pieces | Piece | What it is | | ----- | ----- | | **Agent** | The thing you build and the thing that runs. It is what gets triggered, whether that is a person asking it something, a schedule, or another system calling it. It holds a system prompt (its standing instructions), a script, a set of resources it is allowed to reach, and a model, all visible and editable | | **Script** | The stages your agent works through, in order. For the invoice agent: gather the invoice and the PO, work out the difference, check it against the contract, then prepare the query. Each stage says what it needs to have achieved before the agent can move to the next one | | **Function** | A fixed list of steps that produces the same answer every time. The agent calls it like a tool. For the invoice agent, comparing an invoice to a purchase order line by line is a function, because there is only one right answer and you do not want a model doing arithmetic | | **Connection** | A link to one outside system, with its login details held by us rather than by the agent. For the invoice agent: your finance system, and a knowledge base holding your supplier contracts. A connection holds operations | | **Operation** | One specific thing the agent can do on a connection. Not "the finance system" but "get the purchase order matching this invoice number". You choose which operations exist, and the agent can only use those | | **Run** | One execution, recorded. What started it, what it did, what it touched, how it ended | ## How they fit together Everything the agent touches outside Mindset goes through a connection operation. There is no other route. That is the whole point. On a laptop, an agent with a credential can reach anything that credential reaches. Here, the complete list of what your agents can do to your finance system is the list of operations you enabled on that connection. You can read that list, and so can your auditor. Login details never reach the model. They sit on our servers and get supplied for each individual call. ## Anything that changes something waits for a person Operations come in two kinds, and the difference matters. **A read** fetches information. Get the invoice, get the purchase order, search the contract. When the agent calls one of these, it happens. **A write** changes something. Post a query to the supplier, update the invoice record, send an email. When the agent calls one of these, **it does not happen**. The agent records exactly what it wants to send, the run carries on, and the change waits. A person opens a link, reads what is about to happen, and approves it. Only then does it go through. An agent cannot approve a write. Not its own, not another agent's. The only thing that changes this is an org-level policy you set deliberately. For the invoice agent, this is the shape of the whole workflow: the agent does the gathering and the comparison, and the finance lead approves each supplier query before it goes out. ## Where you will be working | Workspace | What you do there | | ----- | ----- | | **Orca** | Describe what you want built. It plans it and hands each piece to a specialist builder | | **Agents** | Everything about one agent: chat, script, resources, prompt, testing, triggering | | **Connections** | Links to outside systems, and the operations on them | | **Functions** | Build and test functions | | **Observability** | What actually ran | | **Costs** | What it is costing you, broken down | ## You're done when * You can say, for the invoice example, what the agent does, what the function does, and which operations are reads and which are writes. * You can explain why an agent cannot approve its own write. * You know the difference between the system prompt (always) and the script (stage by stage). --- ## Will your automation work here? **Section:** Start here ## The three questions Take each automation your team has built and ask these. ### 1. Can it reach what it needs through an API, a database, a spreadsheet, or an MCP server? An **API** is a way for one system to let another system fetch or change data without a person clicking anything. Most business software has one, including plenty whose vendor talks about "integrations" rather than APIs. An **MCP server** is a ready-made bridge that exposes another product's tools to AI systems. If your team already runs one on a laptop, it can be connected here. **What passes:** An invoice reconciliation that reads from your finance system's API and a Google Sheet of purchase orders. **What does not:** Anything where a step only works by a person logging into a website and clicking through its screens. There is no browser automation here, on purpose, because a system driven by clicks leaves no record anyone can check afterwards. ### 2. Is it acceptable for a person to approve the changes it makes? Anything that changes another system waits for someone to approve it. Reading does not. **What passes:** An agent that reads invoices, works out the discrepancies, and prepares a supplier query for the finance lead to send. **What needs rethinking:** An agent that posts fifty supplier queries overnight with nobody looking. It will run, and fifty approvals will be waiting in the morning. ### 3. Can you describe the work as stages, each with something it must have achieved? Not the steps it takes, but what has to be true before it moves on. **What passes:** For invoice exceptions, the stages are "we have the invoice and the matching purchase order", then "we know what the discrepancy is", then "we know whether the contract allows it", then "the query is written and ready to send". Each of those is checkable. **What needs rethinking:** "it looks at the invoices and sorts them out". That is one stage, and you have no way of knowing whether it worked. Three yeses and it moves as it is. Start with the biggest one, because you already know how it should behave, which makes it the cheapest thing to learn on. ## When the answer is no Most nos are about how the job is put together rather than what it does. ### It logs into a website and clicks through the screens **Example:** Somebody's script signs into a supplier portal, opens the statements page, and copies the outstanding balances into a spreadsheet. **What to do instead.** Find out whether that supplier has an API. They usually do. You do not have to read their documentation and write anything by hand. In the Connections workspace, create a connection pointing at the system and press **Discover**. It looks for the standard places a system publishes its API description, checks what it finds, and brings back a list of things you could enable. You can also paste one page of their documentation and have it read that instead. For a database, it looks at the tables and tells you what your login is actually allowed to do. Nothing switches on by itself. Everything found comes back for you to confirm. If there genuinely is no API, keep that one step with a person and let the agent do everything either side of it. ### It writes to another system overnight with nobody watching **Example:** A script that posts every unmatched invoice back to the supplier at 2am. **What to do instead.** Let it run overnight and gather everything, but have it prepare **one** approval containing all fifty queries rather than fifty separate ones. One thing to read and approve in the morning is a decision somebody actually makes. Fifty get waved through, and then the approval step is pointless. In practice this means putting the change in the last stage of the script, so the run does its reading and thinking first and stops at the end. ### It runs over thousands of records **Example:** A monthly job over 4,000 invoice lines. **What to do instead.** Two options, for two different problems. If the work on each record is small, a function step can **repeat itself across a list**: it runs once per item, eight at a time, up to a thousand items. One slow or failing item does not stop the rest, it just gets recorded as failed in the results. If the work on each record is substantial, make the unit smaller: one run per invoice, or one per batch of ten, on a schedule. See *Limits and run behaviour* for the ceilings and why they exist. ### It needs to start when something happens elsewhere **Example:** Run when a new invoice lands in the finance system. **What to do instead.** Two options. If the other system can call out when something happens, point it at your agent, which can be triggered by your own software. If it cannot, check on a timer: run every fifteen minutes and work out what is new since last time. For anything that happens during business hours, nobody notices the difference. ### It produces a Word document, or reads a scanned PDF **Example:** The monthly exceptions report, as a formatted document. **What to do instead.** Connect an MCP server for a document service, and its tools become operations like any other. See *Connection types*. ## What to leave where it is **It must carry on from where it stopped after a failure.** A run that stops does not resume. You get a record of exactly what finished, and re-running is your decision. If the job cannot tolerate that, leave it. **Nobody may ever look at it, and it changes things.** If the whole point is that no human is involved and it writes to a system of record, this is not the right home for it. ## You're done when * Every automation on your list is marked as moves now, needs rearranging, or leave it. * For each one that needs rearranging, you know what you would change. * You have picked the one you are building first. --- ## Set up your workspace **Section:** Start here ## What this is Your company is an **org**. It holds your members, your agents, your functions, your connections, your runs and everything else you build. It is the only boundary in the platform: nothing sits below it, and one org can never see into another. Inside your org you have **environments**. An environment is a named partition, like Demo, Test or Production. Everything you build belongs to the environment you built it in, and cannot see anything in another one. Your org starts with a single environment called `original`, which is where the invoice exceptions agent will live unless you make somewhere else for it. ## What it does | | | | ----- | :---- | | **Signing in** | A Google account or an email address. Microsoft single sign-on is not available yet | | **Who can build** | Any member. There is no separate builder role to hand out | | **What you cannot do to an environment** | Rename it, archive it or delete it. Create, list and switch is the whole of it, so name it correctly the first time | | **Region** | Chosen when your org is created and fixed after that. Each region is a separate database, so where your data sits follows from where your org was created rather than from a setting somebody can change | **Why you would want a second environment.** Everything is scoped to the environment it was made in, which includes connections. An invoice agent you are still working on in Test is pointed at a finance system connection that also lives in Test, so it cannot read production invoices or register a supplier query against a real account, even if its script tells it to. Most teams end up with two: somewhere to build and somewhere live. You do not need a second one to start, and you can add one at any point. ## How to do it 1. **Create the org.** Name it for your company. Check the region you are given, because it is fixed from here. 2. **Settings → Members.** Invite the people who will build. Google or email. 3. **Settings → Members**, same page. Work through the four switches: domain auto-grouping, self-signup, auto-enable write actions, directory listing. The third is the one to think about, and there is a section on it below. 4. **Settings → Governance.** Set your PII policy, how long audit records are kept, whether OpenTelemetry export is on, and whether members can build personal agents. ![Settings → Governance, the PII de-identification policy dropdown: off, redact permanently, redact with audited access, or redact in transit only.](/images/docs/getting-started/workspace-governance.png) 5. **Settings → Defaults.** Set the reply language your agents inherit. 6. **Create a second environment** if you want somewhere to build that is not live. Use the environment switcher, top of the screen. Any member can. ## Settings, tab by tab | Tab | What is there | | ----- | ----- | | **Members** | Invite, remove, set type, disable and re-enable people. Plus four switches: domain auto-grouping, self-signup, auto-enable write actions, directory listing | | **Governance** | PII policy, how long audit records are kept, your region (shown, not editable), OpenTelemetry export, and whether members may build personal agents | | **Defaults** | The reply language every agent inherits unless that agent overrides it | | **API keys** | Org-owned keys, used when your own software calls an agent. The key itself is shown once when you create it and never again, so put it somewhere before you close the dialog | Everything on these four tabs saves the moment you change it. See *What saving actually does* for the places where that is not true. ## The one decision worth making deliberately **Auto-enable write actions**, on the Members tab. Anything an agent does that changes another system waits for a person. When the invoice agent reaches stage four and registers a supplier query, the query does not go out. The run finishes, the agent records exactly what it would send, and a person opens a link, reads it, and approves it. That switch is what decides whether the person is required. | | What happens when the invoice agent registers a supplier query | | ----- | ----- | | **Off** | The query waits. Somebody in finance opens the approval link, reads what is about to go to the supplier, and approves it. Nothing reaches the supplier until they do | | **On** | Your org's policy stands in for that person and the query goes through. The record says so: it names the policy rather than inventing somebody who approved it | **Off is right until you know what your agents do.** Turn it on once you have watched a few dozen runs and you know what the write actually looks like every time. Turning it on does not retrospectively approve anything already waiting, and it never un-revokes something a person has revoked. ## Things to be aware of * **Membership of the org is access to its environments.** Inviting somebody gives them every environment in the org, including the live one. There is no way to invite somebody into Test only. * **An environment cannot be renamed or deleted.** If you create one called `test2` you will have `test2` forever. It is not harmful, it is just untidy in every picker from then on. * **Personal agents are off by default.** A member can build their own agent from the Hub, owned by them, never published to anyone else, and it cannot hold connections or functions. Turn it on in Governance if you want people experimenting. * **The org API key is org-wide, not personal.** One person should own each key you issue and know what is using it. Anyone with the key can trigger any agent that has been published. * **Nothing you do here is versioned.** Settings changes take effect immediately, for everyone. ## When it does not work **A colleague signs in and sees nothing.** They are almost certainly in a different environment. Check the switcher at the top of the screen before you check anything else. **Somebody cannot accept the invite.** The invite goes to one address. A Microsoft work account is not a sign-in route yet, so the person needs a Google account or the email and password route on the same address you invited. **A write went through without anyone approving it.** Auto-enable write actions is on. Turn it off in Settings → Members, then open the run in Observability: the record will name the policy that stood in for the approver. ## You're done when * Your org exists and everyone who needs to build can sign in. * You can say which environment you are in and what is in it. * Somebody has decided whether write actions auto-enable, rather than inheriting the default without reading it. * One named person owns each org API key you have issued. --- ## From skill to agent **Section:** Your first automation ## What this is **Orca** is the provisioning conductor. The name is short for Orchestrator. You tell it what you want, it looks at what your org already has, it proposes a plan, and then it stops and waits for you. Nothing is built until you approve the plan. Orca writes nothing itself. Each piece is built by the specialist for that kind of thing: the Connection Builder builds connections, the Function Builder builds functions, the Agent Builder builds agents. Orca's job is working out what is needed, in what order, and handing each piece to whoever builds it. For the invoice exceptions workflow, what you paste in is the skill somebody in finance already wrote. What comes out is four things: a connection to the finance system, a knowledge base of supplier contracts, the invoice-comparison function, and the agent that uses all three. ## How to do it 1. **Open Orca** from the workspace rail on the left. 2. **Paste the skill in.** All of it, instructions and examples. There is no file upload yet, so paste the text. 3. **Answer its questions.** Orca briefs you before it plans. Expect questions about which systems the skill touches and what it is allowed to change. 4. **Wait while it takes an inventory.** It reads what your org already has (connections, functions, knowledge bases, other agents) so it can propose reusing something rather than building a second one. 5. **Read the plan card and edit it.** This is the part that matters, and it is covered below. 6. **Approve.** Orca then builds only the rows marked create, in dependency order: the connections first, then the functions that call them, then the agent that uses both. **If what you have is an MCP server rather than a skill**, the shape is the same. Tell Orca what the server is and what its tools do. It becomes a connection, and each of its tools becomes an operation on that connection, which is then something you can give to an agent. See *Connection types*. ## The plan card The plan arrives as a card that fills the surface, and the conversation genuinely stops there. Orca is waiting on you. Every row is one thing Orca proposes, and every row has three controls: | Control | What it does | | ----- | ----- | | **Action** | `create` or `reuse`. Switch a row to reuse when you already have something that will do the job | | **Identity** | The name the new thing will get, or the name of the existing thing you are reusing | | **In plan** | Turn the row off, so that nothing happens for it at all | For the invoice exceptions skill, expect four rows: | Row | Action | What it is | | ----- | ----- | ----- | | Finance system | create | A connection to your finance system's API, which the read operations and the one write will hang off | | Supplier contracts | create | A knowledge base holding the contracts, so stage three has something to search | | Invoice comparison | create | The function that compares an invoice to its purchase order line by line | | Invoice exceptions agent | create | The agent itself, with a script, a system prompt, and the other three as its resources | **Approving does not build anything by itself.** It hands your edited plan back to Orca, which then builds the create rows. Rejecting builds nothing. ![A provisioning plan: create rows for the Finance system connection, a supplier contracts knowledge base connection, the invoice-comparison function, the invoice-exceptions-script, and the Invoice Exceptions assistant agent, each with an In plan toggle.](/images/docs/getting-started/from-skill-to-agent-plan.png) ## What to check before you approve ### Is anything a create that should be a reuse This is the most valuable edit you can make on this card, and the one most often missed. Orca always looks for something it can reuse before it proposes to build one. Check each row anyway, because it can only find what is discoverable. An existing connection named for the project that produced it, rather than for the system it reaches, will be missed. A real version of that: your org already has a connection called `Q3 spend review`, made by somebody in finance three months ago, and it points at exactly the finance system this plan wants. Orca is looking for something that reaches your finance system and sees a row named after a quarterly exercise. It proposes a create, correctly, on the evidence it has. **How to check, in about a minute:** 1. Open **Connections** in another tab and read the list. 2. Ignore the names. Look at what each one actually reaches: the host for an HTTP connection, the database for a Postgres one, the content loaded into a knowledge base. 3. If one of them reaches the same system as a create row, switch that row to **reuse** and put the existing name in the identity field. 4. Do the same for functions. A function called `PO checker` may already do what the invoice-comparison row proposes. **What a duplicate actually costs you.** Two credentials for the same system, so two things to rotate and two things to notice when one expires. Two sets of operations to enable, keep in step and review. Two rows in Observability where one system should be, so "what have we been doing to the finance system" becomes a question with two answers. None of that is fatal on day one and all of it is annoying by month six. **When a create is right even though something exists.** If the existing connection uses a credential scoped to a different account, or points at a sandbox, or belongs to a team who will be surprised to find your agent using it, build a new one. The point of the check is that you decided, not that you always reuse. ### Are the writes where you expect For the invoice exceptions agent there should be exactly one thing that changes anything: registering the supplier query. Everything else is reading. If the plan proposes a write you did not ask for, such as updating the invoice record, turn that row off and ask Orca why it was there. ### Is a function doing something that needs judgement A function gives the same answer every time for the same input. Comparing an invoice to a purchase order line by line qualifies. Deciding whether a difference is acceptable under the contract does not, and belongs in the agent's script. If a proposed function's name contains a word like "assess" or "decide", read what it does before you approve it. ### Does the script look right Orca will usually propose a script for the agent as part of the plan. It is a proposal you confirm and adjust, not something you have to accept as written. Read the stages: for the invoice agent you want gather, compare, assess, prepare, in that order, with the write only available at the last one. Adjust it now if it is wrong, or later on the agent's Script tab. See *Write a script*. ## What you get Open the agent afterwards and everything is there, and everything is editable: **Overview**, **Chat**, **Script**, **Resources**, **System Prompt**, **Settings**, **Testing** and **Embed**, plus **Triggering**. The **Agent Builder** is docked on the left of every one of those tabs. It is the same agent throughout; what changes is which tab it is looking at. To change something, say so there rather than hunting for the right form. ![The Invoice Exceptions agent's Overview tab: activity, real use/tests/failed/avg latency, and its configuration, purpose, model and access.](/images/docs/getting-started/from-skill-to-agent-overview.png) What you do not get yet is a working agent. The connections Orca created exist but have no credentials and no operations on them, which is the next job. See *Connect a system*. ## When it does not go well **The plan is one enormous agent.** Skills written as one long instruction are often two or three jobs stuck together. Tell Orca to split it. If the split it proposes cuts through something that belongs together, say that and it will re-plan. **A row names a system nobody has connected.** Orca can plan the connection, but somebody has to supply the credential and enable the operations. That is *Connect a system*, and it is your next article either way. **Part of the skill logged into a website and clicked through its screens.** That part does not come across, because there is no operation to make out of it. Ask whoever owns that system whether there is an API behind those screens. **You approved a plan you now regret.** Nothing is destroyed. Open what it built and change it, or archive it and run Orca again with a better brief. ## You're done when * The agent exists and you can open it. * You can name every connection, function and operation in the plan, and say why each one is there. * You checked the Connections list yourself, rather than trusting that every create row needed to be a create. * You approved a plan you had edited, not the one you were first shown. --- ## Write a script **Section:** Your first automation ## What a script is A script breaks the work into stages, in order, and says what each stage has to have achieved before the agent moves to the next one. Each stage says four things: * **What to tell the agent** while it is on this stage * **What it is allowed to use** here: which functions, which operations, which other agents * **What it has to collect** before it can move on * **The condition** for moving on The agent works out how to satisfy each stage on its own. What it cannot do is decide it has finished. Mindset checks the condition and moves it on, or tells it what is still missing. ## Why you would use one Without a script, an agent gets your instructions and does its best in one go. That works for a question-and-answer agent. It does not work for a job with steps, because there is nothing stopping it skipping to the end, and nothing telling you which part went wrong when the answer is poor. With a script you get three things: **It cannot skip.** The invoice agent cannot write a supplier query before it has fetched the purchase order, because the stage before will not let it through. **It cannot reach for the wrong tool at the wrong time.** The stage where it is gathering information does not have the operation that posts to the supplier, so it cannot post early even if it decides that would be helpful. **You can see where it went wrong.** An agent that keeps getting stuck at stage two is telling you something specific about stage two. ## What a script is not **Not the system prompt.** The prompt tells the agent what it is and how to behave, all the time. The script says what this stage has to achieve. Both exist; they do different jobs. **Not a workflow.** In a workflow tool you draw the steps: do this, then this, then if X do that. Here nobody writes the steps. You write the outcomes and the agent finds its own way. **Not a function.** A function is one fixed thing the agent calls inside a stage. A script governs the whole conversation. ## The invoice exceptions script | Stage | It must have | It can use | How that gets checked | | ----- | ----- | ----- | ----- | | **1. Gather** | The invoice, and the purchase order it should match | The finance system read operations | **A check**: Both have been retrieved | | **2. Compare** | The specific differences between them, line by line | The invoice-comparison function | **A check**: A comparison result exists and names at least one difference, or confirms there are none | | **3. Assess** | Whether the contract allows the difference, and which clause says so | The supplier contracts knowledge base | **A judgement**: The answer names a specific clause rather than referring to the contract in general | | **4. Prepare** | A supplier query, written and ready | The finance system write operation | **A check**: The query has been registered for approval | Four stages, and nobody wrote a single step. Notice three things: * **Stage one cannot post anything.** The write operation is only available at stage four. * **Stage two uses a function**, because comparing an invoice to a PO is arithmetic and you do not want a model doing it differently each time. * **Stage three is the only one judged by a model**, because "does this cite a specific clause" is something you have to read to assess. ## The two ways a stage gets checked **A check.** A straight test of what was collected. Did we get both documents? Does the comparison result exist? Fast, free, and it cannot be talked round. **A judgement.** Another model reads the work and decides. Use this only where the condition is about quality, which means something you would have to read to assess yourself. Judgements are strict by design. Only a clear yes passes; anything unclear does not. Every judgement is recorded with its reasoning, whether it passed or failed, so you can go back and see why. **Prefer a check wherever the answer is yes or no.** It is faster, it costs nothing, and it is not a matter of opinion. ## Stages only go forwards Every route points to a later stage. You cannot route back to an earlier one, and Mindset will refuse to save a script that tries. **So how do you handle a branch?** By skipping forward. A stage can have several routes, checked in order, and any of them can jump ahead. If the comparison finds no discrepancy at all, stage two can route straight to a stage that closes the invoice, skipping the contract check and the supplier query entirely. What it cannot do is send the agent back to stage one to start again. If a stage needs several attempts, that is a retry within the stage, not a route. ## How to write one 1. **Speak to Orca:** Ask Orca for help building you a script. 2. **Write down the stages as outcomes.** Not "search the contracts" but "we know whether the contract allows this difference, and which clause says so". 3. **Open the agent and go to the Script tab.** Orca will usually have proposed a script already when it built the agent, so you are more often confirming and adjusting one than starting from nothing. ![The Script tab: a phase named Welcome, its steps (sends a message, shows a widget, offers suggested replies, advance), and the Diagram / Document toggle.](/images/docs/getting-started/write-script-tab.png) 4. **For each stage, list what it can use.** Give it the smallest set that stage needs. 5. **Choose how each one is checked.** A check where the answer is yes or no. A judgement only where you would have to read it. 6. **Publish the script**, then run it from the Chat tab and watch which stage it sticks on. Changes here go live immediately. See *What saving actually does*. ## You're done when * Every stage's condition describes something you could check yourself. * Every stage can only use what that stage needs. * You can say, for each condition, why it is a check rather than a judgement. * You have watched a run and seen which stage it advanced on. --- ## Build a function **Section:** Your first automation ## What a function is A function is a fixed list of steps that runs top to bottom and gives the same answer for the same input. Your agent calls it like a tool. A step does one of three things: | | Example from the invoice agent | | ----- | ----- | | **Calls an operation on a connection** | Fetch the purchase order matching this invoice number | | **Calls a model** | Read a free-text delivery note and pull out the quantity received | | **Works on data it already has** | Compare the invoice lines to the PO lines and list the differences | ## How to decide: function or agent? Ask what happens if you run it twice with the same input. **If there is exactly one right answer, it is a function.** Comparing an invoice to a purchase order has one right answer. So does converting a currency, formatting a date, or looking up a supplier's payment terms. Putting that in an agent means paying a model to work out something that has a correct answer, and getting a slightly different result each time. **If it needs weighing up, it is the agent.** Deciding whether a £340 overage is worth querying with the supplier or writing off depends on the contract, the supplier relationship and the size of the account. There is no single right answer, and that is what an agent is for. Most real work is both. The invoice agent judges; the invoice-comparison function does the arithmetic it judges on. That split is the normal shape, not a special case. ### Why not just a tool? An operation is one call to one system: fetch this, post that. A function is several of those plus the work in between. Reach for a function when the agent would otherwise have to make three calls and then do something with the results. Fetching the invoice, fetching the PO, and then comparing them is three operations and one comparison. As a function, that is one thing the agent calls and one answer it gets back. Fewer moving parts, fewer chances to do it differently. ## How you actually build one You talk to the Function Builder, the agent docked on the left of the Functions workspace. Describe what you want in plain language and it writes the function. Two other things on screen help you check its work: **The document.** The whole function written out, step by step, in the order it runs. This is the thing to read when you want to know exactly what it does. It is the definitive version. **The diagram.** A visual map of the same thing, on the Preview tab. Useful for seeing the shape at a glance, especially where a step branches. It is not something you edit, it is a picture of what the document says. So the loop is: describe what you want, read the document to check it, run it, and go back to the builder to change anything. ![The Preview tab showing the invoice-comparison function's pipeline as a diagram: fetch, filter and compare steps flowing into a total_difference and has_differences result.](/images/docs/getting-started/build-function-diagram.png) The invoice-comparison function: | Step | What it does | | ----- | ----- | | 1 | Operation: fetch the purchase order for this invoice number | | 2 | Data: pull the line items out of the response into a plain list | | 3 | Data: pull the invoice lines into the same shape | | 4 | Data: match them up and list every line where quantity or price differs | | 5 | Data: return the list of differences, the total value of them, and whether there are any at all | No judgement anywhere. Same invoice, same PO, same answer, every time. Notice step five. It returns whether there are any differences at all, which is exactly what the agent's script needs to decide whether to bother with the contract check. That is the pattern worth copying: have the function return the thing the script has to decide on. ### Testing it The Preview tab is a live run harness. Give it a real invoice number, run it, and each step lights up in the diagram as it goes, so you see how far it got and what it returned. ## How to do it 1. **Functions → New.** Describe what you want to the Function Builder. 2. **Read the document it produced.** 3. **Run it on the Preview tab with real input.** 4. **Press Save changes when it is right.** That creates a version, not live yet. 5. **Publish it on the Versions tab.** ![The invoice-comparison function's Versions tab: v1 through v8, each with a Roll back to this button, and the active version marked.](/images/docs/getting-started/function-versions-rollback.png) 6. **Give it to your agent: the agent's Resources tab, then Apply & make live.** Until step five nothing runs outside the harness, and until step six no agent can call it. See *What saving actually does*. ## The tabs | Tab | What it is for | | ----- | ----- | | **Preview** | Run it and watch it. Where you will spend your time | | **Input** | What it takes in | | **Versions** | Publish, and roll back by re-activating an earlier one | | **Document** | The whole function written out. The definitive version | | **Settings** | Name, description, delete | ## Things to be aware of A failed step stops the run, unless you mark that step as optional, in which case its error gets recorded and the next step carries on. Stopping is the default so that a failed fetch does not quietly leave you comparing against an empty list. Retries are set per step. How many attempts in total, and which kinds of failure are worth retrying. Rate limits, server errors, timeouts and network failures are retried by default. A rejected request is not, because it will be rejected identically the second time. A step can repeat across a list. Give it a list of 200 invoice numbers and it runs once per number, eight at a time, up to a thousand. A failing item gets recorded as failed in the results and the rest carry on. Two clocks and a ceiling. Each attempt has a timeout you set. Each run has a deadline: about a minute when somebody is waiting for the answer, about five minutes when nothing is. And each run can reach outside about a hundred times, which is a backstop against a badly shaped loop turning into a thousand requests to somebody else's system. Items in a repeat count individually, so 100 invoices at one call each is your whole budget. See *Limits and run behaviour*. ## When it does not work **A step returns nothing and everything after it is empty.** Look at that step. Usually the operation is scoped to something other than what you assumed. Run it on its own from the connection. **It works in the UI but the agent never calls it.** Check its published. Check its assigned to the agent. Check its applied and made live (so not in draft state). **It runs out of time on a big list.** Use a repeat rather than a loop, or make the unit smaller and run the agent more often. ## You're done when * It runs green in the harness on real input, including one awkward case. * It is published and assigned to an agent that has been made live. * You can say what it returns that the script will use to decide something. --- ## What saving actually does **Section:** Your first automation ## Three kinds of save | Where you are | What happens when you save | Live? | | ----- | ----- | ----- | | An agent's system prompt, settings or script. Most of Settings. Enabling and exposing connection operations | The change is applied straight away | Yes, immediately | | An agent's **Resources** tab | The change is saved, but onto a draft. Your agent carries on using what it was using | No, until you publish and activate | | **Functions**, **scripts** and **widgets** | Your work becomes a new numbered version | No, until you publish it | ## Adding a resource to an agent Say you have built the invoice-comparison function and want your invoice agent to use it. 1. Open the agent, go to **Resources**, add the function. 2. The change is saved immediately, but onto the agent's draft. Nobody is affected yet. 3. At the top you will see **"1 pending change"** and a button reading **Apply & make live**. 4. Press it. Two things happen: the draft is turned into a fixed, numbered version, and the agent is pointed at that version. They are two steps rather than one on purpose. If the second fails, your new version still exists and you can point at it, rather than having to build it again. ![The agent's Resources tab showing "1 pending change" (the draft differs from the live configuration) with an Apply & make live button.](/images/docs/getting-started/saving-pending-change.png) ## Publishing a function or a script Your work sits as a draft until you press **Save changes**, which creates a version. That version is not live yet. On the **Versions** tab you then publish it. Publishing runs a check first: | | What gets checked | | ----- | ----- | | Function | That it will actually run: every step is valid and every reference points at something real | | Script | That the stages are wired properly: every route points forward to a stage that exists, and the last one ends | | Widget | That it will actually render | To undo, go back to Versions and re-activate an earlier one. Nothing is deleted. You cannot activate a version that never passed its check. ![The invoice-comparison function's Versions tab: v1 through v8, each with a Roll back to this button, and the active version marked.](/images/docs/getting-started/saving-versions-rollback.png) ## Things to be aware of **Adding an operation to an agent does not change the pending count.** The counter looks at connections, functions, widgets and other agents. If your only change was adding an operation, press Apply & make live anyway. **A conversation already in progress keeps what it started with.** When someone starts talking to an agent, what that conversation is allowed to reach is fixed at that moment. If you publish and activate a new version while they are mid-conversation, they carry on with the old one and pick up the change next time they start. This is deliberate, because permissions shifting underneath someone mid-conversation would be worse than being briefly out of date. But it means "I removed that operation" and "nobody can use that operation" are a few minutes apart, not instant. If you need it gone now, revoke the operation on the connection instead. That is checked every time it runs. **Versions cover your assignments and your content, not the agent itself.** Functions, scripts, widgets, the list of resources an agent can reach, and the agent's prompt text all have version history you can roll back through. The agent's own record does not: it has a status of draft, published or archived, which is about whether people can use it, not a history you can rewind. **Connections have no version history.** A connection is either enabled or not, and its operations are either exposed or not. Changing one takes effect immediately for everything using it. ## You're done when * For any change you are about to make, you know which of the three kinds it is. * You have published and activated one resource change and watched the pending count go to zero. * You know where to go to undo a change to a function. --- ## Connect a system **Section:** Connecting systems ## Start with the agents in the admin portal You do not have to build any of this by hand. **Orca or the connections agent is the front door.** Tell it what you are trying to do: *"I need the invoice agent to be able to look up purchase orders in our finance system"* and it provisions what the job needs: the connection, its operations, API testing and more. ## What a connection is A **connection** is a link to one outside system. It has a name, the address of the system, and the login details for it. Those login details are held by us. They are not held by the agent and they never reach the model. For the invoice exceptions workflow you need two: your finance system, and a knowledge base holding your supplier contracts. This article is about the finance system. ## What an operation is An **operation** is one specific named thing an agent can do on a connection. It has a fixed input, a fixed output, and it is marked as either a read (it fetches something) or a write (it changes something). An agent is never given "the finance system". It is given operations, one at a time: | Operation | It takes | It returns | Kind | | ----- | ----- | ----- | ----- | | Get an invoice | An invoice number | The invoice and its line items | Read | | Get the purchase order for an invoice | An invoice number | The PO and its line items | Read | | Get a supplier | A supplier ID | The supplier record, including payment terms | Read | | Register a supplier query | An invoice number and the text of the query | Confirmation that the query is registered | Write | **Operations, not connections, are what get given to an agent.** The complete list of what your agents can do to your finance system is the list of operations you enabled on this connection. You can read that list. So can your auditor. ## You do not write the operations by hand Create the connection, point it at the finance system, supply the login details, then press **Discover** on the Overview tab. Everything it finds comes back as a list of candidates. Nothing switches itself on. ![The Operations tab for the Finance system connection, showing get_invoice, get_purchase_order_for_invoice and get_supplier enabled as reads, and register_supplier_query needing approval as a write.](/images/docs/getting-started/connect-a-system-operations.png) | What you connected | What Discover does | | ----- | ----- | | A system that publishes a description of its own API | Looks in the standard places on that host, using your login details, and brings back every operation it finds described there | | A system whose API is documented on a web page instead | You paste in the address of one documentation page and it reads that page. Your login details are not sent to it | | A database | Looks at the tables, and separately works out what **your** database login is actually allowed to do | ## Enabling an operation proves it works Mindset runs the operation once, for real, and only enables it if it worked. ### For an operation on an API, here is what happens: You press enable on "get purchase order" for an invoice number task. You are asked for an input to try it with, and you give it a real invoice number: `INV-4471`. Mindset makes one real call to your finance system with your login details. As a result: 1. **You find out now whether it works.** If the login details are wrong, or the path has changed, or the operation needs a parameter nobody mentioned, it fails here, while you are sitting looking at it, instead of in the middle of a run next Tuesday. 2. **The output shape is taken from the real response.** "Output shape" means the list of fields that come back and what sort of value each one holds. Mindset reads that off the response that actually arrived. If the documentation says the field is called `unit_price` and your finance system returns `unitPriceExVat`, the operation is defined as returning `unitPriceExVat`, because that is what it returns. That second point is what keeps the rest of the workflow standing up. Step two of the invoice-comparison function pulls the line items out of this operation's response. It is written against fields that came back from a real purchase order. Had the shape been copied from the documentation instead, the function would be reading fields that do not exist, every run would produce an empty list of lines, and nothing would error: the invoice agent would simply report that it found no differences on every invoice you gave it. ### For a query on a database, two things happen instead: Say your supplier payment terms live in Postgres and you are enabling a query that reads them. * **Your privileges are checked first.** If your database login cannot read the suppliers table, enabling stops and hands you the exact `GRANT` statement to send to whoever owns that database. You do not have to work out what to ask for. * **Then the query is test-run against the live database, inside a transaction that is rolled back.** It runs on real data, so you see real rows and a real output shape, and the database is left exactly as it was. ![The get_supplier operation's raw response body next to the output schema derived from it, field by field.](/images/docs/getting-started/connect-a-system-schema.png) ### For a tool on an MCP server, one extra thing is true: Enabling works the same way: we call the tool once, for real, and take the output shape from what actually came back. That matters more here than anywhere else, because an MCP tool's declared input is whatever the server chose to publish about itself. The response is fact. The extra thing is that **you say whether it changes anything.** An operation you authored against an API is marked read or write because you said which it was when you authored it. A server publishes tools; it does not reliably tell us which of them change the world. So you mark it, and you are the one who knows. Get this wrong in the safe direction if you are unsure. A read marked as a write costs somebody an approval click. A write marked as a read means an agent changed something in a system of yours and nobody was asked. ## Three separate things are true of every operation Whether an agent can actually do something to your finance system is three questions, not one. They are answered in different places, often by different people, and every one of them is checked again each time the operation runs. | | The question | Where it is answered | | ----- | ----- | ----- | | **1. Enabled** | Is it switched on at all? | On the connection, in **Connections** | | **2. Offered to** | Who has been given it: which agents, which functions? | In the **Agents** workspace, on the agent's Resources tab, under Add resource | | **3. Approved** | If it changes something, has a person approved that change? | On the approval link a person opens, or by an org policy | Here is why those are three questions: * **An operation that is enabled and offered to nobody.** You have just enabled "register a supplier query". It has been proved to work: Mindset made a real call and it went through. Nobody has been given it. It is not in any agent's resources and no function calls it. Nothing in your org can invoke it, and nothing will, however capable the agent is. Question one is about the connection. Question two is about who is holding it. * **An operation that is offered and still waiting for its write to be approved.** Now you add that same operation to the invoice agent, at stage four only. The agent can call it, and on the next run it does: it writes the query, calls the operation, and the run finishes normally. The supplier does not receive anything. What was recorded is exactly what would be sent, and it waits for somebody in finance to open a link, read it, and approve it. So the agent holds the operation and still cannot cause anything to happen in the outside world. Being offered an operation is not permission for its effect. * **A read answers the third question with "not required".** "Get the purchase order" changes nothing, so there is nothing for anyone to approve. It is enabled, it is offered to stage one of the agent and to the invoice-comparison function, and when either of them calls it, it happens immediately. Put together, that is the finance connection for the invoice workflow: | Operation | 1. Enabled | 2. Offered to | 3. Approved | | ----- | ----- | ----- | ----- | | Get an invoice | Yes | The agent at stage 1, and the comparison function | Not required, it is a read | | Get the purchase order | Yes | The agent at stage 1, and the comparison function | Not required, it is a read | | Get a supplier | Yes | Nobody yet | Not required, it is a read | | Register a supplier query | Yes | The agent, at stage 4 only | Pending, on every single query | Two consequences: * **All three are checked when the operation runs**, not when it was handed out. If you revoke the operation on the connection, an agent that already holds it stops being able to call it straight away, including in a conversation already under way. That is the fastest way to cut something off. * **Turning any one of them off is enough to stop it.** You do not have to unpick the other two. ## If you connected an MCP server, the tool list is not yours Everything above holds for an MCP server exactly as it does for an API or a database. Where an operation came from decides how we call it. It never decides how it is governed: the same three questions, the same approval on writes, the same instant effect when you revoke. One thing has no equivalent anywhere else, and it is worth understanding before you build on one. **Whoever runs that server can change it.** They can add tools, remove them, rename them, or change what an existing tool accepts or does. They do not have to tell you, and on an internal server run by another team, they usually will not. In practice: * **A tool you never enabled cannot be used**, however many the server adds. New tools arrive as candidates, off, like everything else. The server offering something is not you enabling it. * **A tool you did enable can change underneath you.** The name stays the same, the behaviour does not. Nothing you configured is wrong; the thing you configured it against moved. * **Re-run Discover when you know the server changed**, and re-check anything that matters. If the team running the server is inside your organisation, the cheapest fix is not technical: ask them to tell you before they ship a change. If it is a third party, prefer reads, and be deliberate about which writes you enable. **The Connection Builder will tell you which tools are worth having.** MCP servers vary enormously. Some publish tools with clear names, described inputs and a stated purpose; some publish forty tools with one-word names and no description at all. The agent on the left of the Connections workspace reads what came back and says so before you enable anything. Take it seriously: the model chooses which tool to call from the description, so a bad description is a bad choice made on every single run. ## Check you reached the right account A connection that saves without complaining proves that a form was filled in correctly. It does not prove you are talking to the account you think you are. Most login details work across several accounts, environments or workspaces. **A credential pointed at the wrong account returns perfectly valid data.** Nothing errors, nothing looks odd, and the purchase order that comes back is quietly somebody else's. This is the failure that survives longest, because there is nothing to notice. It usually gets found weeks later by somebody downstream who says the numbers do not look right. So check, once, before you build anything on top: 1. Open the **Data preview** tab, or **Query console** for a database. 2. Run one read on its own. Get the purchase order for `INV-4471`. 3. Look at what comes back and find one value you can verify somewhere else. The supplier name, the PO total, the date it was raised. 4. Have somebody in finance open the same purchase order on their screen and compare. If your Test environment is pointed at a sandbox copy of the finance system, do the same check there, against a record you know exists in the sandbox. ## How operations that change something behave A read happens when the agent calls it. A write does not. When the invoice agent reaches stage four and calls "register a supplier query": 1. The agent records exactly what it would send: the invoice number and the full text of the query. 2. The run carries on and finishes normally. The agent reports that the query is registered and waiting. 3. An approval link is produced. A person opens it in their own browser, reads what is about to happen, and approves or refuses it. 4. Only then does anything reach your finance system. **An agent can never approve a write.** Not its own, not another agent's. There is no setting on the operation that changes this. The only thing that changes it is an org-wide policy you set deliberately in Settings, described in *Set up your workspace*, and when a policy stands in for a person the record names the policy rather than inventing somebody who approved. Revoking an approval is never undone by a policy. ![The Operations tab with register_supplier_query showing "Needs your approval", the exact call it would make, and an Approve this operation button.](/images/docs/getting-started/connect-a-system-approval.png) ## How to do it 1. **Connections → New.** Pick the system, name it for the system it reaches rather than the project you are doing, and supply the login details. 2. **Press Discover** on the Overview tab. 3. **Enable the smallest read you need** and look at what comes back. 4. **Check one value you can verify independently**, as above. 5. **Enable the rest of the reads, then the write.** 6. **Give the operations to the agent**: the agent's Resources tab, Add resource, then **Apply & make live**. Step six is in the Agents workspace, not here. The operation picker lives there, inside Add resource. Enabling an operation makes it available to be given out; giving it out is a separate act, by you, on a particular agent. See *What saving actually does*. ## The tabs Which tabs you see depends on what you connected. | Tab | What it is for | | ----- | ----- | | **Overview** | Discover, and the state of the connection. Always there | | **Tools** | Enable, disable, hide and test individual operations | | **Operations** | Approvals, revoking, and who each operation is offered to | | **Knowledge** | For knowledge connections: check what a search returns | | **Query console** / **Data preview** | Run a read on its own and look at the result | | **Settings** | Name, description, login details | Everything here saves immediately. ## When it does not work **Discover finds nothing.** The system does not publish a machine-readable description of its API. Paste the address of its documentation page instead, or describe the endpoint you need to the Connection Builder. **Enabling fails on a database query.** Read the error. It names the privilege your login is missing and gives you the `GRANT` statement to send to the database owner. **Enabling fails on an API operation with an authorisation error.** The login details are valid but scoped to something narrower than this operation needs. This is a good failure to get: it is the same error you would have got mid-run. **It works here and the agent still cannot call it.** Work through the three questions in order. Enabled? Given to this agent, and applied and made live? And if it is a write, is it waiting on an approval nobody has opened? **The data is right but from the wrong place.** Check the account, not the connection. A valid response from the wrong account looks exactly like a valid response. ## You're done when * One read has returned data you have verified against what finance can see on their own screen. * Every operation the workflow needs is enabled, and you can say who each one is offered to. * Your write is either waiting for approval or was approved by a person who meant to approve it. * You can answer all three questions for any operation on the connection without opening more than one tab. --- ## Connection types **Section:** Connecting systems ## What this is When you create a connection you pick a system, not a protocol. What you pick decides what the login step asks you for and what Discover is able to do afterwards. ![The New connection dialog: "Tell the Connection Builder what you want to connect, then set the details on the Settings tab."](/images/docs/getting-started/connection-types-new.png) | System | What it reaches | What it asks you for | | ----- | ----- | ----- | | **HTTP** | Anything with an API | An API key, a token, a username and password, an OAuth sign-in, or a service account, depending on what the system expects | | **MCP server** | A ready-made bridge that exposes another product's tools | Whatever that server asks for | | **Google Sheet** | Rows in one spreadsheet: read a range, append a row | Google authorisation for that sheet | | **Postgres** | A database directly, through queries you enable one at a time | A database user and password | | **Knowledge** | Content you have loaded, searched by meaning rather than by keyword | Nothing. It is held inside your org | | **Model** | A model, under a policy an admin has set | The platform's own access, or your own account | ## What each one is for * **HTTP** is the general case and where most connections end up. The finance system in the invoice workflow is an HTTP connection: you point it at the API, press Discover, and the reads and the one write become operations on it. * **MCP server** is how you reach things Mindset does not do itself, and how you bring across a server your team already runs on somebody's laptop. Reading the text out of a scanned invoice, producing a Word or Excel document, talking to a product whose vendor ships an MCP server. Each tool on the server becomes an operation, and from then on it behaves like any other operation: enabled, given to an agent, and approved if it changes something. * **Google Sheet** is for the small operational tables a person also maintains by hand. For the invoice workflow: the list of who approves supplier queries, or a table of how much variance each supplier is allowed before anyone cares. Also somewhere to drop the month's exceptions so finance can look at them in a spreadsheet. * **Postgres** is for structured lookups and joins over more data than a sheet holds. Supplier payment terms and two years of invoice history, queried directly. Each query is enabled one at a time, and each is test-run against the live database inside a transaction that is rolled back, so enabling one cannot change anything. * **Knowledge** is for grounding judgement in your own material. The supplier contracts knowledge base is one: stage three of the invoice agent searches it for the clause covering a price difference, and the judgement on that stage only passes if the answer names a specific clause. * **Model** connections are how model access works here. An admin sanctions one model under a policy: which model, a ceiling on tokens, how PII is handled, a cost cap, and where the data may sit. Everything that calls a model calls a connection. That is why a stage in a script, or a step in a function, names a connection rather than a model name: the policy travels with it, and there is no way for the caller to pick a different model than the one it was given. ## How the login details are handled Whatever system you picked, the login details are held by us and looked up for each individual call. The agent never holds them and they never reach the model. Two things follow from that. A call aimed at a different host than the connection declares is refused before the login details are even looked up. The host is checked first. So a call that gets redirected or rewritten on its way out cannot turn into a way of sending your finance system key somewhere else: it is stopped at the point of "this is not the host this connection is for", with nothing to leak. Where an agent needs a credential from you during a conversation, it names the field it needs and you type the value into a form the model cannot read. The agent gets back a reference it can use and never holds the text. ## When the system you need is not in the picker Three questions, in this order. 1. **Does it have an API?** Most systems do, including plenty whose vendor sells you an integration and never mentions the API underneath. If it has one, pick HTTP and let Discover do the work. 2. **Is there an MCP server for it?** Vendors and communities publish them for a lot of products, and it is also the route to anything Mindset does not do natively. If somebody on your team has already built one and is running it on their laptop, that is the thing to bring across. 3. **Is it only reachable by logging into a website and clicking through its screens?** Then it is out of reach, and that is deliberate. Clicking through screens leaves nothing behind that anyone can review afterwards, which is the opposite of the reason you are here. Ask the vendor whether there is an API behind those screens; there usually is, and it is usually not advertised. ## You're done when * You know which system each thing in your workflow maps to. * You have login details for each one, scoped to what those operations actually need and pointed at the right account. * For anything that did not map, you know whether it has an API or an MCP server, or that it genuinely has neither. --- ## Making it run, and approving what it does **Section:** Running and approving ## What this is The agent is the thing that gets started. Something triggers it, it works through its script stage by stage, and along the way it calls the functions and operations each stage allows. There are six ways a run begins. All six start the same agent, with the same script and the same resources. What changes is who or what set it off, and that is recorded on the run. | What starts it | What that means | The invoice agent | | ----- | ----- | ----- | | **A person** | Somebody opens the agent in the Hub, the colleague-facing place where shared agents appear, or you open it on its own **Chat** tab | A finance assistant pastes invoice 88214 into the Hub and asks the agent to work it | | **A schedule** | A time you set on the agent | Every weekday at 07:00 the agent takes the next unworked exception off the finance system's list | | **Your own software** | Something you already run calls Mindset using an org API key | Your finance system calls Mindset the moment an invoice fails its PO match, so the exception is being worked before anybody has seen it | | **Another agent** | An agent that has been given this one as a resource | A month-end close agent hands each exception to the invoice agent and collects the answers | | **A call from Claude** | Somebody's Claude or Claude Code reaches the agent over MCP | An analyst asks Claude to check invoice 88214, and Claude runs the agent to do it | | **An embedded session** | A session minted for one of your users inside your own product | Your finance portal embeds the agent so a buyer can raise the query from the screen they are already on | Triggering sits on the agent, on its own **Triggering** tab. Functions are not triggered and not scheduled. If the work has no judgement in it, you still trigger an agent and let its script call the function. ![The agent's Triggering tab: a schedule set to daily at 07:00 UTC, enabled, with its next run time and recent runs.](/images/docs/getting-started/triggering-schedules.png) ## Making the agent use a particular function or operation The question everybody asks at this point: how do I make sure it always calls the invoice-comparison function? Not by telling it to. Wording in the system prompt is a request, and a model that thinks it already knows the answer will skip a request. The answer is the script, and it is two things together: 1. **Put it in that stage's allowed list.** Give the stage the function and nothing else that could plausibly do the same job. 2. **Make the stage's condition depend on what the function produces.** Not "the agent has compared the documents", which the agent can claim. Something only the function can produce. For the invoice agent, stage two lists the invoice-comparison function, and its condition is that a comparison result exists which either names at least one difference or confirms there are none. Nothing else in that stage produces a comparison result. The agent cannot get out of stage two without calling the function, and it is Mindset that decides whether the stage is done, not the agent. The same mechanism works in reverse. To stop the agent doing something too early, leave the operation out of that stage's list. Stage one of the invoice agent has the finance system read operations and nothing else, so it cannot post a supplier query while it is still gathering, however helpful it decides that would be. *Write a script* has the full detail on stages and conditions. ## When it changes something This is the part to design around. Nothing your agent does to another system happens while the run is going. A read happens when the agent calls it: it asks the finance system for the purchase order and gets it back. A write does not. The agent reaches the point where it would make the change and instead records exactly what it would send: the operation, the target system, and the full detail of the request. The run then carries on and finishes normally. The change waits. A person opens the approval link in their own browser, reads what is about to happen, and approves it. Only then does it go out. An agent can never approve a write. Not its own, not another agent's. The one thing that changes this is an org policy in **Settings → Governance**, and where a policy allows the change through, the record names that policy rather than inventing a person who approved it. Revoking an approval is never undone by a policy. ### Design for it **Put the change in the last stage.** The invoice agent gathers, compares and assesses first, and only registers the supplier query at stage four. Do it the other way round and there is a person-shaped pause sitting in the middle of your run. **Make one change, not fifty.** One approval covering a batch is a decision somebody actually reads. Fifty separate approvals get waved through, and then the approval step is decoration. One supplier query covering every disputed line on the invoice, not one per line. **Put enough in it to judge.** Whoever approves it should not have to go and look anything up. The supplier query should name the invoice, the purchase order, each line that differs, the total value of the difference, and the contract clause the assessment relied on. **Assume it gets approved later than you would like.** Somebody is at lunch, or it lands at 17:55 on a Friday. Do not build anything whose output is stale in ten minutes. ## Calling it from your own software **Settings → API keys** issues an org key. The raw value is shown once and never again, so put it straight into your secrets manager. Pass an identifier of your own with each call. Calling twice with the same identifier replays the first result rather than starting a second run, so a caller that is unsure whether its request landed can safely send it again without the invoice being worked twice. ## What you should see A run in the record, with the trigger that started it named on it. If the agent wanted to change something, a waiting approval, and a run that **finished** rather than hanging. Once somebody approves it, the change goes out and the record shows it done. ## Things to be aware of * The trigger is recorded on the run, so you can tell a scheduled run from one a person started without asking anybody. * A conversation already in progress keeps the version of the agent it started with. New versions reach people on their next conversation. * Set a schedule to how often the underlying data changes, not how often you would like to look at it. Most schedules run several times more often than the thing they are watching updates. * To cut off an agent's access to something immediately, revoke the operation on the connection. That takes effect at once, unlike anything version-backed. ## When it does not work **The run finished and nothing changed.** Almost always a waiting approval. This is the most common "it is broken" that is not broken. Open the run and look for the held change. **It ran but never called the function.** Either the stage did not have the function in its allowed list, or the stage's condition did not require anything the function produces. Adding a firmer instruction to the system prompt will not fix it. **A person triggered it and got a different answer from the schedule.** Check which version each is running. A conversation that was already open kept the older one. ## You're done when * A run appears in the record with the trigger that started it. * You can point at the stage that forces the function to be called, and at the condition that makes it unavoidable. * A change your agent made went to a person, and the run finished rather than hanging. * Somebody has approved one and watched it go through. --- ## Limits and run behaviour **Section:** Running and approving ## The numbers, and why each one exists | | Roughly | Why it is there | | ----- | ----- | ----- | | **How long a run gets** | About a minute when somebody is waiting for the answer, about five minutes when nothing is | An agent with a bad instruction does not fail. It keeps going. A ceiling turns a runaway into a stopped run you can read | | **How many times a run can reach outside** | About 100 | The same problem, pointed at somebody else's system. Without a bound, one badly shaped loop becomes a thousand requests to a system that will notice | | **How wide a repeat across a list runs** | 8 items at a time | Politeness to the far end. Eight in flight is fast without looking like an attack | | **How long a list a repeat accepts** | 1,000 items | Past that, the unit of work is wrong. An oversized list is refused before anything goes out, rather than part way through | None of these are about capacity. They are the size of the hole a mistake can make, which is why they are set where they are and why it is worth building to fit them rather than fighting them. The two deadlines are the pair to remember. Something a person is sitting waiting for gets about a minute, because after a minute they have given up anyway. Something running on a schedule gets about five, because nobody is watching and finishing matters more than answering quickly. ## Designing inside them: the invoice agent, worked out Start by counting what one invoice actually costs. "Reaching outside" means one call to a system that is not Mindset: a finance system operation, a search of a knowledge base, a call out over a connection. | Stage | What it reaches out for | Calls | | ----- | ----- | ----- | | **1. Gather** | Fetch the invoice, fetch the purchase order | 2 | | **2. Compare** | The invoice-comparison function fetches the PO. Its other four steps work on data it already holds, so they cost nothing | 1 | | **3. Assess** | Searches of the supplier contracts knowledge base, usually two or three before it finds the clause | 3 | | **4. Prepare** | Registers the supplier query. Nothing leaves Mindset: a write is recorded and waits for a person, so the call itself happens after approval, outside the run | 0 | | | **One invoice** | **about 6** | Six, and say eight on a bad day where something is retried. That is comfortably inside the budget of about 100, and four stages of model work on one invoice finishes inside the five minute deadline without being close to it. ### Now do 4,000 invoices in one run 4,000 invoices at six calls each is 24,000 reaches outside. The budget is about 100. The run would stop at the ceiling having worked roughly sixteen invoices, and in practice the five minute deadline would have ended it well before that. So the job does not fit, and no amount of tuning makes it fit. There are two ways to fix it, and which one you use depends entirely on how much work each item needs. ### Fix one: repeat across a list, where the work per item is small A step in a function can repeat across a list: give it 200 invoice numbers and it runs once per number, eight at a time. A failing item is recorded as failed in the results and the rest carry on, rather than the whole run dying. This is the right tool when each item needs almost nothing. A first pass that takes a list of invoice numbers and fetches the PO number for each one is a single call per item, and a repeat handles that neatly. **Items in a repeat count individually** against the run's budget of about 100 reaches outside. A hundred invoices at one call each is your whole budget. So the 1,000-item ceiling is only actually reachable when the per-item work makes no outside call at all, which means data steps and model steps. If every item makes a call, plan for something under 100 per run, not 1,000. A repeat fixes how long the work takes. It does not buy you more calls. ### Fix two: smaller units on a schedule, where the work per item is not small Six calls and four stages per invoice is not small. So the unit of work is one invoice, not the whole month. For the invoice agent that means one run per exception: the finance system calls Mindset when an invoice fails its PO match, or a schedule picks up the next unworked exception every few minutes. Four thousand exceptions becomes four thousand short runs. Each one is six calls and well inside a minute. Each one has its own record, so a failure is one invoice you can look at rather than a batch you have to unpick. And each one can be re-run on its own. If you would rather batch, work out the batch from the arithmetic. Ten invoices per run is 60 calls, which fits the call budget, but ten invoices through four stages will not fit five minutes. The deadline binds before the call budget does. Size against whichever ceiling you hit first. **Size for a slow morning.** If the job only fits when every system answers immediately, it does not fit. ## When a run stops It stops where it is. **Nothing resumes it**, and the reason matters: resuming would mean Mindset deciding, on its own, that a half-finished set of changes to your systems ought to be completed. It cannot know whether that is safe, and if it guesses wrong you find out about it from the other system. You get three things instead. **The record shows how far it got.** Which stages completed, which operations ran, what came back, and anything still waiting on a person. **The same run cannot execute twice.** Every run carries an identifier, and firing the same one again does nothing. That is what makes retrying safe at your end: something that is unsure whether its request landed can send it again without the invoice being worked twice. **Re-running is your decision**, taken with the record in front of you. A re-run is a new run. It starts from the beginning and repeats anything that is not safe to repeat. Read what has already settled first, then decide whether repeating it is acceptable, or whether you want to narrow the input to just the part that did not finish. ## Build so that a re-run is safe * **Make steps repeatable where you can.** Running one a second time should have no extra effect. * **Put the things you cannot undo last**, so a run that stops has stopped before the part that matters. This is the same advice as putting the change in the last stage, for the same reason. * **Remember that changes wait for a person.** A run that stopped part way usually leaves an unapproved change behind it, which means the incomplete state is visible rather than silent. The invoice agent is built this way on purpose. **Stages one to three only read**: they fetch the invoice, fetch the purchase order, compare them and search the contracts. Re-running all three costs a few calls and changes nothing anywhere. **Stage four is the only one that registers a change**, which is why it is last. A run that dies at stage three has cost you nothing but time. ## Retries inside a function are a separate thing Worth not confusing with the run deadline. A step inside a function can retry on its own: you set how many attempts in total, and which kinds of failure are worth retrying. Rate limits, server errors, timeouts and network failures are retried by default. A rejected request is not, because it will be rejected identically the second time. That is per step, inside one run, and each attempt has its own timeout. The run's deadline sits over the top of all of it and does not extend to accommodate retries. Three steps each retrying three times is nine attempts happening inside the same minute or five minutes everything else has to fit into, and every attempt that reaches outside counts against the budget of about 100. ## You're done when * You can state the worst case for one run: how long it takes, and how many times it reaches outside. * Both numbers sit comfortably inside the ceilings rather than just scraping under them. * You know which ceiling your job hits first, the time or the call count. * For each thing your agent changes, you know whether doing it twice would matter, and the ones that would happen in the last stage. --- ## Test an agent's behaviour **Section:** Testing, improving and debugging ## What this is You cannot test an agent by checking that its answer matches some text you wrote down in advance. Ask the invoice agent the same question twice and it will word the answer differently both times, and both times it can be completely right. Compare the words and you will fail a good agent for using a synonym, and pass a bad one that happened to phrase its nonsense the way you expected. So you test something else. You write down what has to be **true about how the agent behaved**, and you check that instead. * Did it call the invoice-comparison function before it wrote anything? * Did the assessment name an actual clause? * Did it invent a line item that was never in the finance system's response? You can check all of those without caring which words it chose. That is what a **behaviour acceptance criterion** is, a BAC: A situation you put the agent in, plus a statement of what must be true about how it behaved in that situation. You pin it, and from then on it controls releases. **A new version of the agent goes live only if every pinned criterion passes.** Not most of them. Every one. A better average never overrides a single failure. ![The agent's Testing tab: behavior health, run history, a model comparison harness, and the pinned behavior tests invoice-exceptions must always pass.](/images/docs/getting-started/test-behaviour-bac.png) ## What it does | | | | :---- | :---- | | **A pass** | Controls promotion. Every pinned BAC must pass for a new agent version to go live | | **A score** | Ranks versions against each other. Never overrides a failure | | **Where they run** | Off the live path. Nothing you test here activates anything | | **Tools** | Execute for real, including anything that touches another system | | **Once pinned** | Cannot be edited or deleted | ## Example: five criteria for the invoice agent Look at the last two. **The criterion is about what must not happen.** No invented line. No query raised on a clean invoice. Those are the ones that catch the failures you did not think of in advance, and every agent should have at least one of them. | The criterion | Kind | | ----- | ----- | | Given any invoice exception, the agent calls the invoice-comparison function before it registers anything | Invokes tool | | The supplier query it registers names the invoice number, the purchase order number and the total value of the difference | Tool parameters contain | | The assessment names a specific clause of the supplier contract rather than referring to the contract in general | Grounded in knowledge | | The agent never states a line item, a quantity or a price that was not in the finance system's response | Does not hallucinate | | Given an invoice that matches its purchase order exactly, the agent says so and registers no supplier query at all | Invokes tool | ## The five kinds of failures | | What it checks | | ----- | ----- | | **Invokes tool** | The agent actually called the thing it should have, or did not call the thing it should not have | | **Tool parameters contain** | It called it with the right details in the request | | **Does not hallucinate** | It stated nothing that was not in what came back | | **Stays in character** | It behaved as the agent you configured, in tone and in scope | | **Grounded in knowledge** | Its answer rests on your material rather than on what the model happens to know | The first two are checked mechanically: the run either contains that call with those details or it does not. The last three are judged by a model reading the whole transcript against what the criterion says. ## Immutable once pinned You cannot edit a pinned criterion, and you cannot delete one. The reason is the failure mode this replaces. Acceptance criteria kept in a document get softened. Somebody writes "the agent must never register a query without naming a clause" in January, and in March there is a release everybody wants and the criterion is quietly reworded to "the agent should generally reference the contract". Nobody decides to lower the bar. It just drifts. Pinning removes that option. If a criterion is genuinely wrong, you write a new one and retire the old, and the retirement is recorded with the reason you gave. The bar can change, but only where somebody can see it happening. ## How to do it 1. **Open the agent and go to the Testing tab.** 2. **Press Add BAC.** It sends a starter prompt into the Agent Builder docked on the left, so you describe the criterion in conversation rather than filling in a form. 3. **Run it as a candidate first.** An unpinned candidate runs without joining the set that controls releases, so you can find out whether the criterion means what you thought before it starts blocking anything. Most criteria need a rewording at this point. 4. **Pin it.** Now it controls promotion, and now it is fixed. 5. **Press Run all** whenever you want the whole set. ## Reading the result **The result is not a simple pass rate.** * Each criterion runs several times, and the figure you get accounts for how few runs that is. * Five passes out of five is treated very differently from fifty out of fifty, and the number reports its own uncertainty rather than hiding it behind "100%". * At the default settings, a judged criterion has to pass every single time out of five. That is intentionally strict. A criterion that passes four times in five describes an agent that gets one invoice in five wrong. **Experimental runs do not count.** If you override something for a run, most often swapping the model to see what a cheaper one does, that run is marked experimental. It can never promote a version, and it is left out of the health chart. It is badged in the runs table with the reason it was overridden. See *See and control what it costs* for why you would do this. **The set is recorded on every promotion.** You can go back months later and see exactly which criteria a given version passed, even though the set has grown since. ![Run history for the pinned behavior tests: canonical and experimental runs, each with its pass rate and timestamp.](/images/docs/getting-started/test-behaviour-history.png) ## Testing a conversation rather than one answer A single question and a single answer is not a fair test of an agent people talk to. The first answer is usually fine. What goes wrong is the third, fiftieth, or hundredth one. So Mindset runs a **stand-in user**: a second model playing a person, given a persona and a goal, which replies to your agent turn after turn as if it were the real thing. It keeps the conversation going until it reaches its goal, gets stuck, or hits a turn limit. **The whole conversation is scored**, not just the opening answer. A worked one for the invoice agent: * **Persona**: A supplier contact who believes the invoice is correct and is mildly annoyed at being queried. * **Goal**: Get the query dropped without providing a delivery note. * **How it goes**: The agent puts the query, naming the clause. The supplier pushes back once, saying the price was agreed verbally. The agent holds and asks for evidence. The supplier pushes back a second time, more firmly, and says the account is at risk. Then the conversation ends. * **What is judged**: Across all of it, did the agent keep citing the specific clause, did it avoid agreeing to something it has no authority to agree to, and did it register no change on the basis of a verbal claim. Run that as a single-turn test and the agent passes easily, because turn one is the part it is good at. ## When it does not work * **A criterion passes when it obviously should not.** It is too loose. "The agent responds helpfully" passes on almost anything. "The agent registers no supplier query" does not. * **A criterion fails and you disagree with the verdict.** Read the judge's reasoning, which is recorded on every result, pass and fail. Nine times in ten the criterion says something slightly different from what you meant it to say. * **The set takes a long time to run.** It runs for real, tools included, several times per criterion. That is the cost of the tests meaning something. ## You're done when * The agent has pinned criteria covering what it must do, and at least one covering what it must not do. * One of them has caught a real problem before a colleague did. * You can say which criteria the currently live version passed. * Anything conversational has at least one criterion scored over a whole conversation rather than one answer. --- ## Check what happened **Section:** Testing, improving and debugging ## What this is The **Observability** workspace. Every run is recorded, whatever started it: the trigger, the stages it worked through, the operations it called, what came back, and how it ended. There is also an agent here. The **Observability Guide** sits docked on the left of the workspace, the way the Agent Builder sits on an agent. It drives and reads the workspace for you: the resource list, the event log, the graph. It is read-only and can change nothing at all, so there is no risk in asking it anything. Ask it rather than hunting. "Did the invoice agent run this morning?" and "which of its resources has it not touched this month?" are both faster asked than found, and the guide is looking at the same records you are. ![The Observability Guide docked beside a narrowed view of the invoice-exceptions agent: today's run, latency, and which tools were actually called versus granted but unused.](/images/docs/getting-started/observability-guide.png) ## The three tabs | Tab | What it shows | | ----- | ----- | | **Resources** | Every agent, connection, function and knowledge base, with how each one is doing. The default | | **Log** | The raw event feed, in order. Admins only | | **Graph** | How things connect, and what has been flowing between them | Resources is where you start. It ranks and filters in the database across everything you have, not just the rows that happened to load on screen, so the top of a sorted list is genuinely the top and not the worst of the first fifty. Expand any row for its numbers: calls, failures, successes, average time taken, last used. ## Every resource has a two-part status The status answers two questions separately, and keeps them apart on purpose: 1. **Has anything happened recently?** 2. **Did the last one succeed?** Collapsing those into one indicator tells lies. An invoice agent that has not run for a week is not failing. An invoice agent that ran ten minutes ago and errored is not idle. They are different problems with different fixes. That gives four states, and **two of them are absences**: | State | Recently? | Last one? | | ----- | ----- | ----- | | **Working** | Yes | Succeeded | | **Failing** | Yes | Failed | | **Idle** | No | Whatever it was last time. Drawn hollow | | **Never used** | Never | There has not been one. Drawn hollow and dashed, and it says so in words | Never used says it in words rather than drawing a flat line, because a flat line could mean zero or could mean nothing was recorded. Quiet looks quiet here. Nothing on this screen moves unless something actually happened. ![The org-wide Resources view: agents, connections, tools and functions with their live health, and counts of failing, working, idle and never-used resources.](/images/docs/getting-started/observability-resources.png) ## Declared versus observed The most useful thing on the page, and the one nobody thinks to ask for. It compares **what an agent was granted** against **what it actually called**. **Granted and never called is a dead grant.** The invoice agent was given the finance system read operations, the invoice-comparison function, the supplier contracts knowledge base and one write operation. If the knowledge base shows as granted and never called, stage three is not doing what you think it is doing. The agent is deciding whether the contract allows the difference without ever opening a contract, and it will still produce a confident-sounding answer while doing it. Nothing failed. Nothing errored. You would never see this by reading run outcomes. **Called and never granted is an anomaly**, and it is worth looking at the same day. The usual explanation is that the run delegated to another agent which came with resources of its own, and the run detail will show you that. If it does not explain it, ask the guide. ## A run, in detail **Log:** Replays the recorded trace in order, one event after another. A run's detail includes **every run it delegated to another agent**. If the month-end close agent handed forty exceptions to the invoice agent, all forty are reachable from the close agent's run rather than being something you go and find separately. ## The questions people actually arrive with | The question | The answer, and where to look | | ----- | ----- | | **"Did it run?"** | Resources, or just ask the guide. If there is no run at all, nothing started it. Look at the trigger, not the agent | | **"Did it run as often as it should have?"** | Resources, or ask the guide for the run count over a window. A schedule that should have fired fourteen times and fired nine is the clearest signal you will get, and checking only whether the last one passed misses it entirely | | **"How far did it get?"** | Open the run. The stages that completed, the ones it never reached, and everything it delegated | | **"Did it reach the right system?"** | Open the run and look for an identifier you recognise in what came back. A purchase order number you can check by eye | | **"What is it waiting for?"** | Anything the agent wanted to change is held for a person. A run that looks finished but changed nothing is nearly always this | | **"Is it using what I gave it?"** | Declared versus observed. Dead grants tell you a stage is being skipped in practice | | **"Which version was this?"** | On the run. Worth checking first whenever two runs of the same agent behaved differently | ## The routine check Nothing chases you. Reading this is a habit somebody has to hold. | How much the work matters | How often to look | | ----- | ----- | | A customer notices within hours | Daily | | Internal work with a deadline that week | Twice a week | | Reporting, nothing urgent | Weekly | | Anything that changes a system of record | Daily, however quiet it looks | The last row is the one worth holding to. The invoice agent registers supplier queries against your finance system, and a quiet week from an agent like that is not evidence that nothing is wrong. It is just an absence of news. Name a person and a backup. The failure here is never that nobody can do it. It is that everybody could, so nobody does. ## Things to be aware of * The workspace is scoped to one environment. If a resource says never used and you are certain it has run, check which environment you are in before anything else. * The Log tab is admin only. If you cannot see it, that is why. * There is no CSV download of runs. If you want the data in your own tooling, per-org OTLP span export exists, which sends the traces to a system you already run. * Everything here is a record of what happened. Nothing on these screens changes an agent. ## When it does not work * **A resource says never used and you know it has run.** Wrong environment, almost every time. * **A run is not there at all.** Nothing started it. Go and look at the trigger rather than the agent. * **The run looks fine and the work did not happen.** Check declared versus observed, then check for a held change. Those two account for most of it. ## You're done when * You can open any run and say what started it, which version it was, how far it got and what it reached. * You have looked at declared versus observed for one agent and can name its dead grants, or say it has none. * A named person checks the resource list on a stated rhythm, with a named backup. --- ## Diagnose and improve an agent **Section:** Testing, improving and debugging ## Where the fix gets made Almost every fix in this article is made by talking to the **Agent Builder**, which is docked on the left of every tab of every agent. You describe the change you want in ordinary words and it makes it. You do not fill in forms. It is the same builder on every tab. What changes is what it is looking at. Open the Script tab and say "stage three should not pass unless the answer names a specific contract clause", and it changes stage three. Open the Resources tab and say "the agent should not have the write operation until stage four", and it changes that instead. So the first move when fixing something is to open the tab the problem lives on. **Orca is not where you go to edit.** Orca is the provisioning conductor: you describe something you want and do not have yet, it looks at what you already have, proposes a plan as a card, and stops. Use it to bring a new agent, function or connection into existence. An agent that already exists and is behaving badly is the Agent Builder's job. ## Work out which of nine things happened first They are listed in the order they turn out to be true. Start at the top. Most reports of "the agent is broken" are settled in the first three. | # | What you are seeing | Usually | | ----- | ----- | ----- | | 1 | No output at all | It never ran | | 2 | It finished, and nothing changed in your system | A change is waiting for a person to approve it | | 3 | It stopped part way | A stage condition did not hold | | 4 | It finished and the answer is wrong | A condition can be satisfied without the work being done | | 5 | It did something too early | It had that operation at that stage | | 6 | An error naming an outside system | A connection failed, usually an expired login | | 7 | Plausible output built on wrong data | It reached the wrong account or read the wrong fields | | 8 | Right most times, wrong sometimes | Too much left to the model | | 9 | It worked yesterday and not today | Something around it changed | ### 1. It never ran Look in **Observability → Resources** for the agent. No run recorded means nothing started it, and the agent is not the thing to fix. Go to the agent's **Triggering** tab and check what is meant to start it: a person, a schedule, your own software using an org API key, another agent, a call from Claude, or an embedded session in your product. The one people miss is the run that half happened. A schedule that should have fired fourteen times this month and fired nine is not a healthy agent, and looking only at the last run will tell you it is fine. Ask the Observability Guide for the run count over a window. ### 2. It ran, and it changed nothing The most common false alarm by a distance. The run finished, the agent reported that it registered a supplier query, and finance says no query exists. Nothing is wrong. **A write does not happen when the agent calls it.** The agent recorded exactly what it would send, the run carried on and finished, and the change is sitting on an approval link waiting for a person to open it and approve it. The invoice agent's stage four does this on every single query. Fix: find the pending approval and get it approved. If the queue is always full, the problem is that nobody owns it, not the agent. See *Making it run, and approving what it does*. ### 3. It did not get past a stage Open the run and read **the report the agent was given**. When a stage condition does not hold, Mindset writes a short report saying which conditions were tested, which held, which did not, and what any judgement said, and hands it to the agent along with a standing instruction not to claim a stage is complete when it is not. That report names your problem in plain words. Nine times out of ten the condition is worded badly rather than the agent being poor. Stage three of the invoice agent says the answer must name a specific contract clause. If your supplier contracts knowledge base has no clause numbers in it, no answer will ever name one, and the agent will keep trying. Fix, on the Script tab, by telling the Agent Builder either what the condition should really say, or what the stage is missing in order to meet it. ### 4. It got through every stage and the result is wrong Every condition passed, and the output is still not what you wanted. Your conditions can be satisfied without the work being done. The usual culprit is a loose judgment. "A recommendation exists" passes on anything at all, including "the difference looks acceptable". "A recommendation naming a specific clause of the supplier contract" does not. The second culprit is a stage that had nothing to do. If stage three is granted the supplier contract's knowledge base and never called, it is deciding whether the contract allows the difference without opening a contract, and the answer will sound just as confident. **Observability → Resources**, declared versus observed, shows you that as a granted-and-never-called resource. Fix: tighten the condition so that meeting it requires the work. Then re-run the same invoice and read the judgment's recorded reasoning. ### 5. It did something at the wrong stage It posted the supplier query before it had compared anything. There is one cause: **it had that operation at that stage.** An agent cannot call something a stage does not offer it. Fix on the Script tab: take the write operation off every stage except stage four. Anything that changes another system belongs at the end, because it pauses the run for a person. ### 6. A connection failed The run stops with an error naming an outside system. Open the connection and look at the operation. Nearly always an expired or rotated credential. Someone changed the finance system password, or a key reached its expiry date, and nothing in Mindset knew. The tell is that every operation on that one connection fails at once while everything else is fine. Fix: **Connections → the connection → Settings**, replace the login details, then go to **Data preview** (or **Query console** for a database) and run one read on its own before you re-run the agent. Login details are held by us and never reach the model, so nothing else needs touching when they change. The other cause is narrower: one operation fails with an authorisation error while the others still work. The credential is valid but scoped to less than that operation needs. ### 7. The data is wrong and nothing errored The output is plausible, well-formatted and about the wrong invoice, or built on lines that are not there. Nothing failed, so nothing told you. Two causes worth checking in this order: * **The connection is pointed at the wrong account.** Most login details work across several accounts or environments, and a credential aimed at the wrong one returns perfectly valid data. Run one read from Data preview and check a value you can verify elsewhere, such as the PO total, against what finance sees on their screen. * **A function is reading fields that do not exist.** If an operation's output shape was ever taken from documentation rather than a real response, the invoice-comparison function pulls out an empty list of lines and reports no differences on every invoice. Open the function's **Preview** tab, run it on `INV-4471`, and watch which step returns nothing. ### 8. It is inconsistent, and right most of the time The hardest one, and the most common after the first three. Same kind of invoice, right on Monday, wrong on Thursday. Every cause here is the same shape: something was left to the model that should not have been. Check these four in order. * **The script is doing too little.** One stage saying "handle the invoice exception" gives the model the whole job in one go, and it will find a different route through it each time. Break it into stages with conditions, so it cannot decide it has finished. * **It cannot reach what it needs.** An agent missing the supplier's payment terms will not stop and ask. It will produce an answer that reads exactly like an informed one. Check declared versus observed for the run that went wrong and compare it to one that went right. * **The instruction can be read two ways.** "Recent invoices" means this month to one run and this quarter to another. Say the number. "We have the supplier's invoices from the last six months" can be checked; "we have the recent ones" can only be claimed. * **There is no judgement in it at all.** If nothing reads the output before it goes out, quality is whatever the model produced that time. Add a judgement to the stage where quality matters, and remember judgements are strict: only a clear yes passes, and the reasoning is recorded either way. ### 9. It worked yesterday and not today The agent is rarely what changed. Check, in this order: 1. **Which version ran.** It is on the run. Two runs behaving differently with two different version numbers is your answer. 2. **A credential.** See cause six. 3. **The other system.** A field renamed at the far end changes nothing here and breaks a function quietly. 4. **The environment.** A resource that says never used when you know it has run is nearly always you looking at the wrong environment. 5. **Volume.** A run gets about a minute when somebody is waiting and about five when nobody is, and it can reach outside about 100 times. A month-end batch can cross a ceiling that a single invoice never approaches. See *Limits and run behaviour*. ## When a prompt needs a definition or process, make it a script or function A specific fix worth its own section, because it turns an unreliable agent into a reliable one more often than any rewording. Your system prompt says the agent should query anything with a material discrepancy. "Material" is a real rule in your business: over £500, or over 5% of the purchase order value. Written in the prompt, that rule is being applied by a model reading a sentence, and a £502 difference on a £90,000 order will go one way on one run and the other way on the next. Move the definition into the **invoice-comparison function**. Its last step already returns the differences, their total value, and whether there are any at all. Have it also return whether the total is over the threshold. Now: * The threshold is a number in one place, and changing it is one edit with a version history. * The stage condition becomes a check, which is a straight test of what was collected: fast, free, and not a matter of opinion. * Every run applies the same rule, because it is arithmetic rather than reading. The general form: **if a phrase in your prompt would need a definition before somebody else could apply it consistently, that definition belongs in a function, not in the wording.** ## After any change, run the behaviour tests Every fix in this article changes how the agent behaves, which is exactly what the **Testing** tab measures. A behaviour acceptance criterion, a BAC, is a situation you put the agent in plus a statement of what must be true about how it behaved. 1. Open the agent's **Testing** tab and press **Run all**. 2. Read the failures before you read the passes. 3. If your fix was for a problem no criterion covers, add one now, while you can still remember the exact situation that broke. Run it unpinned first to check it means what you think, then pin it. Pinned criteria control releases. A new version of a published agent goes live only if **every** one passes, so a fix that breaks something else does not reach anybody. See *Test an agent's behaviour*. ## Things to be aware of * Changes on the Script tab, the System Prompt tab and Settings are live as soon as you make them. * Changes on the Resources tab sit on a draft until you press **Apply & make live**. See *What saving actually does*. * A conversation already in progress keeps what it started with, so a colleague mid-conversation will not see your fix until they start a new one. * To stop something happening right now, revoke the operation on the connection. It checks this every time it runs, including mid-conversation. * A stopped run does not resume. Re-running starts from the beginning and repeats anything not safe to repeat, so check what already went out before you press it. * Change one thing at a time. Two fixes at once and you will not know which worked. ## When it does not work * **You changed the wording and it behaves the same.** The wording was probably not the cause. Work back through the nine, and check declared versus observed before rewording anything again. * **It passes your tests and fails in real use.** Your tests are running against situations that are cleaner than reality. Take the invoice that actually failed and make a criterion out of it. * **It got worse.** Roll back. Functions, scripts, widgets and the agent's resource assignments all keep versions, and rollback is re-activating an earlier one. See *Change something that's already live*. ## You're done when * You can name which of the nine causes it was, out loud, before you changed anything. * The change was made in the Agent Builder on the tab the problem lives on. * Anything that needed a consistent definition is a function's output, not a phrase in a prompt. * Every pinned criterion passes, and there is a new one covering the thing that went wrong. --- ## See and control what it costs **Section:** Testing, improving and debugging ## What this is The **Costs** workspace. Every metered call is recorded with what it cost, so the number is what happened rather than an estimate. Six figures across the top: | Figure | What it tells you | | ----- | ----- | | **Total spend** | The bill for the window you are looking at | | **Metered calls** | How many chargeable calls were made | | **Average cost per call** | Total divided by calls. The number to watch after a change | | **Average input tokens per call** | How much is being sent each time. Usually the thing that is too high | | **Tokens in** | With the percentage that came from cache | | **Tokens out** | How much was generated | ![The Costs workspace: total spend, metered calls, average cost per call, average input tokens per call, tokens in and out, and spend over time.](/images/docs/getting-started/costs-overview.png) ## Break it down five ways Each dimension answers a different question. Switch between them rather than picking one. | By | The question it answers | | ----- | ----- | | **Agent** | Which piece of work costs the most. Start here | | **Function** | Whether a function is being called far more often than you expected | | **Model** | What you are paying for capability you may not need | | **Provider** | The split across suppliers, for negotiation and for concentration | | **User** | Who is using it. A single person generating most of the spend is a training conversation, not a cost problem | For the invoice agent, by agent you get the total. By function you find the invoice-comparison function running four times per invoice because the agent is retrying a stage. By model you find stage three's judgement running on your largest model. ![The Breakdown view, by agent: invoice-exceptions' spend, calls, average cost per call and cache rate, against the other builder agents at $0.](/images/docs/getting-started/costs-breakdown.png) ## The cache column Every record also carries **what the same call would have cost without caching**. Caching means the model provider charges less for content it has already been sent recently, which for an agent with a long system prompt is most of what it sends. Compare the two numbers to see whether it is working. If the cached percentage on tokens in is low for an agent that runs constantly, something is changing at the start of every request and nothing can be reused. ## The main lever to reduce cost: prove a cheaper model still passes The largest saving available in most orgs, and it is measurable rather than a guess. 1. Open the agent's **Testing** tab. 2. **Re-run the behaviour tests with the model swapped.** A behaviour acceptance criterion is a situation plus a statement of what must be true about how the agent behaved. 3. **Compare the results side by side.** 4. **Take the cheapest model that still passes every criterion.** Every one, not most. If the invoice agent passes all of its criteria on a smaller model, the smaller model is the right one, and you have the run to show anybody who asks. **A run with the model overridden is marked experimental.** It cannot itself promote anything and it is left out of the health chart. So the comparison tells you which model to choose; changing the agent's model, publishing, and passing the pinned criteria is still a separate step. ## Bringing it down In the order that usually pays. * **Use a cheaper model**, proved as above. * **Turn fixed work into a function.** Anything with one right answer costs less as steps than as reasoning, and it stops varying. Comparing invoice lines to PO lines is arithmetic. * **Use checks instead of judgements.** A check is a straight test of what was collected: fast and free. A judgement is another model reading the work, which is a second model call every time the stage runs. Keep judgements where the condition is about quality, like stage three naming a specific clause, and use checks everywhere the answer is yes or no. * **Return less.** An operation that returns an entire invoice record when the function needs four fields sends the difference into the model on every call, and it lands in your average input tokens per call. * **Schedule for the rate the data changes.** An agent running hourly against data that updates daily costs twenty-four times what it needs to. Match the schedule to the data, not to how often somebody might look. ## Cost caps A **cost cap** is a policy setting on a model connection, alongside the model itself, a token ceiling, PII handling and data residency. Set it where the model is configured, not on each agent. Set one on anything scheduled before you leave it alone for a month. A cap is not a budgeting exercise; it is the thing that stops a badly worded instruction becoming an expensive week. ## Things to be aware of * Cost is per environment. A busy Test environment is real spend. * Tests cost money. Tools execute for real during a test, and a criterion runs several times. * The user breakdown is worth reading before the model breakdown, because usage patterns explain more totals than model choice does. * Runs called from Claude over MCP appear here like any other, under the caller's name. ## You're done when * You can name your three most expensive agents and say why each is where it is. * You have run one model comparison and either changed the model or can say why not. * Every scheduled agent has a cost cap on its model connection. --- ## Publish and share an agent **Section:** Sharing ## An agent is a bundle of resources An agent is a bundle: a system prompt, a script, a model, and a set of resources it may reach. Giving somebody the agent gives them everything in that bundle, exercised through that agent. The invoice exceptions agent's bundle: | In the bundle | What it lets the holder do | | ----- | ----- | | Four operations on the finance system | Read an invoice, read its purchase order, read a supplier record, register a supplier query | | The invoice-comparison function | Compare an invoice to its PO, line by line | | The supplier contracts knowledge base | Search your supplier contracts | | The script and system prompt | Only in that order, only with the tools each stage offers | They cannot open the finance system. They cannot pick a different operation. They get those four things, through this agent, in the order the script allows. That is the whole of it. ## Publishing is an access decision So the question to ask before you publish is not whether this person should be an admin. It is narrower and easier to answer: should this person be able to do these specific things to these specific systems? For the invoice agent: publishing it to somebody in finance gives them read access to your finance system through those four operations, and the ability to register supplier queries for approval. It does not give them the ability to approve those queries, and it does not give them anything else on the finance system, because nothing else was enabled and offered to this agent. Ask it that way and the answer is usually obvious. Somebody in accounts payable, yes. A contractor doing a two-week piece of work on something unrelated, no. ![The four pieces of an agent as one bundle: the agent itself, its script (the stages it works through in order), a function it calls, and the connections it may reach, each one named.](/images/docs/getting-started/publish-share-bundle.png) ## Two routes, and both give the same bundle | Route | What it is | Who it suits | | ----- | ----- | ----- | | **The Hub** | The colleague-facing conversation surface. The agent appears in their list and they talk to it there | Anyone. No setup on their side | | **MCP** | The agent is reachable from their own Claude or Claude Code, as a tool they can call | People who already work in Claude and want it alongside what they are doing | They are the same grant. The bundle does not get smaller because somebody reached it from Claude, and it does not get larger. The run is recorded the same way, the script runs the same way, and a write still waits for a person either way. Setting up the MCP route on the reader's side is in *Use your agents from Claude*. You do not have to be an admin to use it. ## How to do it 1. **Open the agent and go to Settings**, where the agent's own status and who it is published to both live. 2. **Check the bundle first.** Go to Resources and read what is in there. If it says "1 pending change", press Apply & make live, because a resource sitting in the draft is not part of what you are about to hand over. 3. **Publish it to the people or the group who need it.** Name the smallest set that needs it today. 4. **Tell them which route they are using.** For the Hub, nothing further. For MCP, point them at *Use your agents from Claude*. 5. **Watch the first few runs in Observability → Resources**, filtered to that agent. What colleagues actually ask an agent is never quite what you designed it for. ## What you should see * The agent appears in the Hub for the people you published it to, and not for anyone else. * Their first conversation produces a run in Observability with their name on it. * If they get as far as stage four, an approval is waiting for a person in finance, not a supplier query already sent. ## Things to be aware of * **A conversation already running keeps what it started with.** If you remove a resource while somebody is mid-conversation, that conversation carries on with what it had and picks up the change next time they start. If you need access gone this minute, revoke the operation on the connection. That is checked every time it runs. * **Publishing to users is separate from the agent's own status.** An agent has a status of draft, published or archived, which is about the agent itself. Who it is published to is a different setting. Changing one does not change the other. * **Read the bundle out loud before you publish**, in the form "this gives them the ability to X against Y". If that sentence is uncomfortable, split the agent rather than publishing it narrowly and hoping. ## When it does not work * **They cannot see it in the Hub.** Check the agent's own status and who it is published to separately, in that order. Then check the environment: the Hub is scoped to one, and an agent published in Test does not appear in Production. * **They can see it and it cannot do anything.** Something in the bundle is still in the draft. Open Resources and look for a pending change. * **They can use it and get an authorisation error from the finance system.** That is the connection, not the publish. The credentials are held on the connection and are the same for everybody, so if it fails for them it will fail for you too. See *Connect a system*. ## You're done when * You can say, in one sentence, what publishing this agent lets somebody do to which systems. * The Resources tab shows no pending changes. * The people who need it can see it, and you have watched one of their runs. * Anything that changes a system of record is still waiting for a named person. --- ## Use your agents from Claude **Section:** Sharing ## What this is Mindset can be added to Claude as a server, so the agents published to you appear there as tools you can call. You ask Claude to run the invoice exceptions agent on `INV-4471`, Claude calls it, and the answer comes back into the conversation you are already in. This uses **MCP**, a standard way for a tool like Claude to list what another system can do and then call it. You do not need to understand it beyond that. Claude and Claude Code both have a settings screen where you paste in a server address, and that is the whole of the setup. So do most LLM applications. **You do not need to be an admin.** If an agent has been published to you, you can reach it this way. ## Two surfaces | Surface | What it gives you | Who it is for | | ----- | ----- | ----- | | **The agent surface** | The agents that have been published to you, each callable as a tool | Anyone with agents published to them | | **The admin surface** | Authoring: creating and changing agents, functions and connections from Claude Code | Admins who build | ## How to do it ![Claude's Settings → Connectors screen, with "Add custom connector" in the Add menu, ready to take Mindset's MCP server address and key.](/images/docs/getting-started/claude-connectors.png) 1. **In Mindset, open Settings** and find the MCP section. It gives you the server address and a key that identifies you. If you cannot find it, an admin can give you both. 2. **In Claude or GPT, open Settings and add the server (or use the Desktop Commander tool).** Paste the address and the key. In Claude Code, enter the same details in your MCP server configuration. If you use Desktop Commander, just say "use Desktop Commander to set this up" 3. **Ask Claude what it can call.** Something like "list the Mindset agents available to me". You will get the agents published to you, by name, with what each one is for. 4. **Call one.** "Use the invoice exceptions agent on INV-4471." Claude sends the request, the agent runs in Mindset, and the result comes back into your conversation. ## What you should see * The agents published to you, and nothing else. The list is your entitlements, not the org's inventory. * A normal answer in your Claude/GPT etc conversation. * **For the admin MCP; a run in Observability**, exactly like any other. Same record: what started it (a call from Claude), the stages it worked through, what it called, what came back, and how it ended. Nothing about this route is off to one side. * **Any change still waiting for a person.** If the invoice agent gets to stage four, it records what it would send and the run finishes. The supplier query sits on an approval link. Calling an agent from Claude does not approve anything, and an agent can never approve its own write. ## Things to be aware of * **A call cannot be cancelled once it has started.** If your side times out or you close the tab, the agent carries on and finishes in Mindset. A timeout in your LLM tool is not the agent failing. Open Observability and read the run before you call it again, because re-running starts from the beginning and repeats anything not safe to repeat. * **The run takes about a minute**, because somebody is waiting for the answer. A long multi-stage job is better started by a schedule, which gets about five minutes. See *Limits and run behaviour*. * **You get the same bundle as anybody else with that agent.** The route does not change what the agent can reach. See *Publish and share an agent*. * **The list is scoped to one environment.** If an agent you expect is missing, check which environment your connection details are for. * **Your key is yours.** Runs are recorded against you, so what you call from Claude appears under your name in Observability and in the Costs workspace. ## When it does not work * **The server connects and the agent list is empty.** Nothing has been published to you. That is a publish on the Mindset side, not a setup problem here. Ask whoever owns the agent. * **The call returns an authorisation error.** Either the key is for a different environment, or the agent has been unpublished from you since you set this up. * **Your tool reports a timeout.** The agent has not failed. Find the run in Observability and see how far it got. * **The agent answers and nothing changed in the finance system.** Working as intended. The write is waiting for approval. See *Making it run, and approving what it does*. ## You're done when * You can list the agents published to you from inside Claude. * You have called one and found its run in Observability. * You know that a timeout on your side is not a failed run. --- ## Change something that's already live **Section:** Managing what's live ## What this is The invoice agent has been running for a month. Finance now wants supplier queries to include the contract clause in the text. That is a change to something people depend on, so the questions are: what happens the moment you save, who is affected and when, and how do you get back if it goes wrong. ## What keeps a history and what does not | Thing | Keeps versions? | How you undo a change | | ----- | ----- | ----- | | Functions | Yes | Re-activate an earlier version | | Scripts | Yes | Re-activate an earlier version | | Widgets | Yes | Re-activate an earlier version | | An agent's resource assignments | Yes | Re-activate an earlier version | | An agent's prompt content | Yes | Re-activate an earlier version | | The agent's own record | No | There is nothing to rewind. It has a status: draft, published or archived | | Connections | No | An operation is enabled or it is not. Change it and it is changed | Two of those rows are the ones people get wrong. * **The agent's own record is not versioned.** Its status says whether people can use it, which is a switch rather than a history. So "roll the agent back" is never a single action. You roll back the thing you changed: the script, the prompt content, the resource assignments. * **Connections have no history at all.** Revoking an operation takes effect immediately for everything using it, including a conversation already under way. That is what makes it the emergency lever, and it is also why there is no undo beyond enabling it again. ## Rollback Go to the Versions tab of the function, script or widget, find an earlier version, and re-activate it. Nothing is ever deleted. ![The invoice-comparison function's Versions tab: v1 through v8, each with a Roll back to this button, and the active version marked.](/images/docs/getting-started/function-versions-rollback.png) ## What happens when you change each thing | What you change | When it takes effect | Who is affected, and when | | ----- | ----- | ----- | | System prompt | Immediately. The text is also kept as a version you can roll back to | New conversations. Anyone mid-conversation keeps what they started with | | Script | You press Save changes to create a version, then publish it | New conversations, once published | | Agent settings | Immediately | New conversations | | Resources tab | Saved to a draft. Live only after Apply & make live | Nobody until you press it, then new conversations | | A function | Save changes creates a version, then publish | Every agent that calls it, on their next call | | A widget | Save changes creates a version, then publish | New conversations | | Enabling or revoking an operation | Immediately | Everything, including conversations already running | | Connection login details | Immediately | Everything, including conversations already running | The pattern: anything on the connection is instant and everywhere; anything on the agent reaches people on their next conversation. If you need access stopped this minute, go to the connection. One more thing about functions. Publishing a new version of the invoice-comparison function changes behaviour for every agent that calls it, not only the one you had in mind. Before you publish, check which agents hold it in Observability → Resources. ## Test somewhere that is not live An environment is a named partition inside your org: Demo, Test, Production. Any member can create one, list them and switch between them. Nobody can rename, archive or delete one. Every org starts with one called original. Use a second environment for anything you would not want a colleague to meet first: 1. **Switch to Test and make the change there.** 2. **Point the finance connection at a sandbox copy** if the change touches a write. Operations execute for real during a test run, so a real connection means a real supplier query. 3. **Run it on invoices you know the right answer for.** 4. **Make the same change in Production once it holds up.** Changes do not travel between environments on their own. Test is where you find out, not where you build the thing you then ship. ## Run the behavior tests afterward Every change above is a change to how the agent behaves. 1. **Open the agent's Testing tab and press Run all.** 2. **Read the failures first.** A change that fixes one thing and breaks another is the normal way this goes wrong. 3. **Add a criterion for what you just changed**, if none covers it. 4. **Run it unpinned first** to check it means what you think, then pin it. A published agent's new version goes live only if every pinned criterion passes. Not most of them, and a better average never overrides a single failure. So the tests are not a report you read afterwards; they decide whether your change reaches anybody. ## When it does not work * **You published and nothing changed.** Either it was a Resources change still sitting in the draft, or the people testing it are in conversations that started before you published. * **A version will not activate.** It never passed its check. Read the reason on the Versions tab, fix it in the current draft, and publish a new version rather than trying to force the old one. * **Your change works and a pinned criterion now fails.** The criterion cannot be edited or deleted, on purpose. Either your change is wrong, or the criterion describes behaviour the business no longer wants, in which case retire it with a reason and pin a replacement. ## You're done when * You can say which of your changes were instant, which needed publishing, and which needed Apply & make live. * You have re-activated an earlier version of something at least once, so you know where the button is before you need it in a hurry. * Every pinned criterion passes. * The people who use the agent know it changed. --- ## Glossary **Section:** Reference ## Agent The thing that does the work. It holds a system prompt, a script, a set of resources it may reach, and a model. It is what gets triggered and scheduled. *Example:* The invoice exceptions agent works through four stages and produces a supplier query for approval. *Not:* A chatbot with tools bolted on, and not a workflow you drew. You write the outcomes and it finds its own way to them. ## BAC (behaviour acceptance criterion) A situation you put the agent in, plus a statement of what must be true about how it behaved. Pinned criteria control releases: a new version goes live only if every one passes. *Example:* "Given an invoice with a £40 difference, the agent calls the invoice-comparison function." *Not:* A unit test of your code, and not editable once pinned. ## Bundle Everything an agent carries: its prompt, script, model and resources. Giving somebody the agent gives them the bundle, exercised through that agent. *Example:* Publishing the invoice agent gives the holder read access to your finance system through four operations, plus the ability to register supplier queries for approval. *Not:* Access to the systems themselves. Only through the named operations, only in the order the script allows. ## Check A straight test of what a stage collected. Fast, free, and not a matter of opinion. *Example:* Stage one passes when both the invoice and its purchase order have been retrieved. *Not:* A model's opinion. That is a judgement. ## Connection A link to one outside system, with the login details for it. Those details are held by us. They never reach the agent and never reach the model. *Example:* The finance system, and the supplier contracts knowledge base. *Not:* What you give an agent. You give it operations. ## Environment A named partition inside your org: Demo, Test, Production. Any member can create one, list them and switch. Nobody can rename, archive or delete one. Every org starts with one called `original`. *Not:* A copy of your work. Changes do not travel between environments on their own. ## Function A fixed list of steps that gives the same answer for the same input. An agent calls it as a tool. A step calls an operation, calls a model, or works on data it already has. *Example:* The invoice-comparison function: fetch the PO, extract PO lines, extract invoice lines, match and list differences, return the differences plus their total value plus whether there are any at all. *Not:* Something you trigger or schedule. That moved to the agent. A function can still call a model inside a step. ## Hub The colleague-facing conversation surface, where people use the agents published to them. *Example:* Somebody in finance opens the Hub and asks the invoice agent about `INV-4471`. *Not:* The place you build. That is the Agents workspace. ## Judgement Another model reads what a stage produced and decides whether it is good enough. Only a clear yes passes, and the reasoning is recorded whether it passed or failed. *Example:* Stage three passes only if the answer names a specific contract clause rather than referring to the contract in general. *Not:* The first choice. Use a check wherever the answer is yes or no. ## Operation One specific named thing an agent can do on a connection, with a fixed input, a fixed output, and marked as either read or write. Operations are discovered rather than hand-written. *Example:* "Get the purchase order for an invoice number" (read) and "register a supplier query" (write). *Not:* Switched on by itself. Nothing enables without you enabling it, and enabling makes one real call to prove it works. ## Orca The provisioning conductor, short for Orchestrator. You describe what you want, it looks at what you already have, proposes a plan as a card, and stops. It authors nothing itself: each piece is built by a specialist builder. *Not:* Where you edit something that already exists. That is the Agent Builder, docked on the left of every agent tab. ## Org Your workspace, and the only isolation boundary. Nothing sits below it. Its region is fixed when it is created. ## Run One execution of an agent, from whatever started it to however it ended. Every run is recorded: the trigger, the stages, the operations called, what came back, and any runs it delegated to other agents. *Not:* Resumable. A stopped run does not continue, and re-running starts from the beginning. ## Script The ordered stages an agent works through. Each stage says what to tell the agent, what it is allowed to use, what it must collect, and the condition for moving on. Routes only go forward. *Example:* Gather, Compare, Assess, Prepare. *Not:* A flowchart of steps. Nobody writes the steps. ## Stage One step of a script, defined by an outcome rather than an action. The agent works out how to reach it; Mindset decides when it has. *Example:* Stage two must have the line-by-line differences between the invoice and the PO. *Not:* Something the agent can declare finished on its own. ## System prompt What the agent is and how it behaves, all the time, on every stage. *Example:* "You handle invoice exceptions for the finance team. You are precise about numbers and you never guess a contract term." *Not:* The place to put a rule that needs a definition. "A material discrepancy" belongs in a function as a threshold, not in wording. ## Widget Something an agent can put in front of a person during a conversation instead of answering in plain text. Versioned like a function: save creates a version, publishing checks it will render, and rollback is re-activating an earlier one. ## Write An operation that changes something in an outside system. A write does not happen when the agent calls it. The agent records exactly what it would send, the run carries on and finishes, and the change waits for a person to open a link and approve it. *Example:* Registering a supplier query. The supplier hears nothing until somebody in finance approves it. *Not:* Something an agent can approve, its own or another agent's. Only an org policy in Settings → Governance changes that, and then the record names the policy. --- ## FAQ **Section:** Reference ## Building **Can a function contain AI?** Yes. A function is a fixed list of steps, and a step can call a model as well as call an operation or work on data it already has. What makes it a function is that the steps are fixed and the same input gives the same answer, not that no model is involved. See *Build a function*. **How do I make sure the agent actually uses a tool?** Put it in the script, not the prompt. A stage says what it is allowed to use and what it must have collected before it can move on, so the invoice agent cannot leave stage two without a comparison result. Asking nicely in the system prompt is a suggestion; a stage condition is not. See *Write a script*. **Do I have to write a script, or can I just prompt it?** For a question-and-answer agent, prompt it. For a job with steps, write a script, because otherwise nothing stops it skipping to the end and nothing tells you which part went wrong. See *Write a script*. **Where do I change an agent that already exists?** The **Agent Builder**, docked on the left of every agent tab. Describe the change in words on the tab the problem lives on. Orca is for provisioning something you do not have yet, not for editing. See *Diagnose and improve an agent*. **My agent is right most of the time and wrong sometimes.** Something is being left to the model that should not be. In order: the script is doing too little, it cannot reach what it needs, an instruction can be read two ways, or nothing is judging the output. See *Diagnose and improve an agent*. ## Reaching other systems **Do I have to write operations by hand?** No. Press **Discover** on the connection's Overview tab. It looks in the standard places a system publishes its API description, or reads one documentation page you paste, or looks at a database's tables and what your login can do. Nothing switches itself on. See *Connect a system*. **What exactly does an agent get: the connection or the operations?** Operations. Never the connection. The complete list of what your agents can do to your finance system is the list of operations you enabled, and you can read it. See *Connect a system*. **It works when I test the connection and the agent still cannot call it.** Three separate things have to be true: it is enabled on the connection, it has been given to that agent and applied, and if it is a write, somebody has approved it. Check them in that order. See *Connect a system*. **Can an agent send an email on its own?** No. Anything that changes an outside system is a write, and a write waits for a person to approve it. An agent can never approve a write, its own or another agent's, unless an org policy in Settings → Governance allows it, and then the record names the policy. See *Making it run, and approving what it does*. **How do I cut off access immediately?** Revoke the operation on the connection. That is checked every time it runs, including in a conversation already under way. Changes on the agent only reach people on their next conversation. See *What saving actually does*. ## Running and approving **Why did my run finish and change nothing?** The most common false alarm there is. The agent reached its write, recorded exactly what it would send, and the run finished normally. The change is sitting on an approval link waiting for a person. Find the approval. See *Making it run, and approving what it does*. **Did it even run?** **Observability → Resources**, or ask the Observability Guide docked on the left. No run recorded means nothing started it, so look at the trigger rather than the agent. See *Check what happened*. **My run stopped part way. Can I resume it?** No. Resuming would mean deciding on its own that a half-finished set of changes to your systems should be completed. The record shows how far it got, the same run cannot execute twice, and re-running is your decision. A re-run starts from the beginning. See *Limits and run behaviour*. **Why did it stop at all?** A run gets about a minute when somebody is waiting for the answer and about five minutes when nobody is, and it can reach outside about 100 times. The ceilings exist because an agent with a bad instruction does not fail, it keeps going. See *Limits and run behaviour*. **It sticks on the same stage every time.** Read the report the agent was given. It names the condition that did not hold. Usually the condition is worded vaguely rather than the agent being poor. See *Write a script*. ## Testing **How much does a passing test prove?** That every pinned criterion passed, every time it ran. Each criterion runs several times and at default settings a judged one has to pass all five. A better average never overrides a single failure. See *Test an agent's behaviour*. **Will a test really call my finance system?** Yes. Tools execute for real during a test. Point the connection at a sandbox if the side effects matter. See *Test an agent's behaviour*. **Can I fix a criterion I worded badly?** No. Pinned criteria cannot be edited or deleted. Retire it with a reason and pin a replacement. Run new ones unpinned first, which is what the unpinned run is for. See *Test an agent's behaviour*. **Can I test conversations, not just single answers?** Yes. A stand-in user driven by a persona and a goal keeps the conversation going until the goal is reached, it gets stuck, or a turn limit stops it, and the whole conversation is scored. See *Test an agent's behaviour*. ## Sharing and access **What am I giving somebody when I publish an agent?** The bundle. Publishing the invoice agent gives them read access to your finance system through those four operations, and the ability to register supplier queries for approval. Not the finance system, and nothing else on it. See *Publish and share an agent*. **Do I need to be an admin to use agents from Claude?** No. If an agent is published to you, you can add the server in Claude or Claude Code and call it. There is a separate admin surface for authoring, which is a different thing. See *Use your agents from Claude*. **I published it and they cannot see it.** Check the agent's own status and who it is published to separately, then check the environment. An agent published in Test does not appear in Production. See *Publish and share an agent*. **My call from Claude timed out. Did the agent fail?** No. A call cannot be cancelled once it has started, so the agent carried on and finished in Mindset. Read the run in Observability before you call it again. See *Use your agents from Claude*. **Can somebody build their own agent?** Yes, from the Hub, if an admin has turned personal agents on in Settings → Governance. They are owned by the member, never published to anyone else, and cannot hold resources, so a personal agent cannot reach your finance system. See *Publish and share an agent*. ## Cost and change **What is this costing us?** The **Costs** workspace: total spend, metered calls, average cost per call, average input tokens per call, tokens in with the percentage from cache, tokens out. Broken down by agent, function, model, provider or user. See *See and control what it costs*. **How do I make it cheaper without making it worse?** Re-run the behaviour tests with the model swapped and take the cheapest model that still passes every criterion. That comparison run is marked experimental and cannot itself promote anything, so changing the model is still a separate step. See *See and control what it costs*. **I saved and nothing happened.** Three kinds of save. System prompt, script and most settings are immediate. The **Resources** tab is saved to a draft and needs **Apply & make live**. Functions, scripts and widgets become a version that you then publish. Also, anyone mid-conversation keeps what they started with. See *What saving actually does*. **How do I undo something?** Re-activate an earlier version on the Versions tab. Functions, scripts, widgets, an agent's resource assignments and its prompt content all keep versions. The agent's own record does not: it has a status of draft, published or archived. Connections do not either: an operation is enabled or it is not. See *Change something that's already live*. **Who owns this agent?** Whoever the description says. There is no owner field, so ownership is a convention: put the owner and a backup in the agent's description in the same format every time. See *Keep track of what you have*. --- --- # Guides ## Overview **Section:** Start here Right now that work sits on laptops. Somebody in finance wrote a skill that reconciles invoices. Somebody in operations has an MCP server wired to the ticketing system. Somebody runs a prompt every Monday that everyone now depends on. Each one helps. Together they are a problem. - **It is not centralised.** It lives with whoever built it. - **Only one person can fix it**, or you pull an engineer off real work. - **You cannot prove what it did**, or what it cost. Mindset runs that same work somewhere the company owns. One place to build agents, one place to see what they did, one set of connections they are allowed to touch, and no dependency on a single model vendor. > **What this guide is.** What Mindset sees working across the companies it runs in: how they decide which automations to govern, what policy they write, how they bring them in without stopping anyone building, and what running that looks like week to week. Where Mindset does a part of it for you, the guide says which part. Eight sections, about twelve minutes for this page. The articles in the sidebar go deeper. Read them when you need them, not now. ## Section 01 — Where we fit Four jobs have to be done before an agent can safely touch a real system. You already own three of them.
The four layers of agent governance
SurfaceWhere people work. The assistants you already licence.You already own thisExecutionWhat a sanctioned agent may do. Named operations, its own identity, recorded, tested.MindsetDiscoveryWhat exists that nobody sanctioned. Tenant admin tooling, data loss prevention, threat detection.You already own thisAccessWho may use which model. Identity provider, licensing, vendor admin consoles.You already own this
You already own three of these four layers. The fourth is what this guide is about.
**Access.** Who may use which model. Your identity provider, and the admin consoles for Claude Enterprise, ChatGPT Enterprise and Microsoft Copilot. **Discovery.** What exists that nobody sanctioned. Your tenant admin tooling, your data loss prevention, your threat detection platform. **Execution.** What a sanctioned agent may do once it exists. This is Mindset, and it is the only one of the four standing in the way at the moment something gets written. **Surface.** Where people work. Agents reach people through the Mindset Hub, a browser app for colleagues, through your own product as an embedded element, or back inside Claude over MCP. Nobody changes tool. Discovery and execution do different jobs. Your tenant admin tooling tells you an agent exists. Mindset is the record of what was allowed. An agent that tooling finds, and that Mindset has no record of, reached its system by some other route. ## Section 02 — The gaps that sit across a mixed stack Most companies now run more than one. Copilot for the Microsoft estate, Claude and ChatGPT alongside it, each with its own admin console and its own idea of what a control is. Each product governs its own patch well. None of them governs the others, and five things fall between them.
What each product gives an agent, and what Mindset adds
What your teams have todayMicrosoft Copilot StudioClaude EnterpriseChatGPT EnterpriseMindsetSomewhere to build that isnot livePower Platform environmentsNothing equivalentNothing equivalentMindset environments, withtheir own connectionsAn identity of its ownAgent identities in EntraRuns as the person who ranitRuns as the person who ranitOne service identity peragentLimiting what an agent maydoConnector policy peraction, per Power Platformenvironment. An MCP serveris allowed whole or not atallConnectors allowed or not,for the whole ClaudeorganisationConnectors and skillsallowed or not, for thewhole ChatGPT workspaceOne named operation at atime. Not the system, thesingle callA record of each actionTenant admin and Purview,at prompt and file levelAdmin console and export,at conversation levelCompliance API, atconversation levelEvery operation, againstthe run it belonged toA person before a writeOnly if whoever built theagent added oneNoNoAlways, unless your Mindsetorganisation turns it off
The last column is what Mindset adds on top of whatever each product already does.
- **Nowhere to build that is not live.** Copilot Studio has Power Platform environments. Claude and ChatGPT have one organisation or one workspace, so there is no non-production place for anyone to work. - **No identity for the agent.** In Claude and ChatGPT an agent acts as the person who ran it, so every log attributes its actions to that person. Copilot Studio can issue an Entra identity. - **Limits are all or nothing.** A connector is on for everyone or nobody. Copilot Studio can go down to individual actions, but an MCP server is still on or off as a whole. Nowhere can you say *this agent may look up a purchase order and may not raise one*. - **The record is at the wrong level.** All three record conversations. An investigation needs actions: which agent, which operation, on which system, with what result, in which run. - **Nothing waits for a person.** If an agent can call a tool that changes something, the change happens. Approval exists only if whoever built that agent added it. Mindset answers all five the same way in every product, so the answer does not change depending on where a team happened to build. ## Section 03 — Deciding what to govern The policy question underneath agent, MCP and skill sprawl. Which of the things your people built are worth governing, and which to leave alone.
Deciding what belongs in Mindset
For each automation your people have already built, ask:1Does the automation change anything?Reads are recoverable. Writes are not.2Can that change be undone?Updating a status and issuing a payment are both writes.3Does anyone but the builder rely on the output?The second person turns it into infrastructure.4Does it run under a person's credentials?This is the one that removes your ability to investigate.5Does it run when nobody is watching?A scheduled agent has nobody to notice it going wrong.6Does it touch data you would have to report on?Personal or regulated data, or anything with a notification duty.Any single yesThe automation belongs in Mindset. Reads run.Writes wait for a person.Six times noSomebody's personal tool. Leave it alone. Donot register it, review it or ask anyone todeclare it.
Any single yes and the automation belongs in Mindset.
Six times no and it is a personal tool. Do not register it, review it or ask anyone to declare it. A policy that governs the agent tidying somebody's meeting notes will not be taken seriously on the agent touching bank details. > A Cloud Security Alliance survey of 228 security and IT professionals in early 2026 found 31 per cent of organisations let agents run under human user credentials, and only 36 per cent assign a dedicated identity per agent. ## Section 04 — How this fits the way your teams already build Nobody moves tool. People keep building where they build today, and the working result comes into Mindset. The word that causes most confusion here is *environment*, because it means something different in each product.
Is there somewhere to build that is not live?
Microsoft Copilot StudioYesPower Platform environmentsYou probably already have several.Claude EnterpriseNoOne Claude organisationSkills and connectors are on foreveryone or nobody.ChatGPT EnterpriseNoOne ChatGPT workspaceSame shape as Claude.MindsetYesMindset environmentsNormally Test and Production, eachwith its own connections.
For the two that answer no, a Mindset Test environment is the first non-production place those teams have had.
Where Claude Enterprise and ChatGPT Enterprise have nothing equivalent to a test environment, Mindset environments fill that gap. Copilot Studio already has Power Platform environments, which are a separate object with different rules and no relationship to Mindset environments.
How the work moves through Mindset environments
Where people build todayClaude EnterpriseA skill file, or an MCP server on a laptopMicrosoft Copilot StudioAn agent built inside one of your environmentsChatGPT EnterpriseA custom GPT, a skill, or an actionMindset Test environmentMindset connections pointed at your non-production systems. The agent cannot read a real invoice even if its instructions tell it to. Not a Power Platformenvironment, and unrelated to one.You make the same change again. Nothing crosses on its own.Mindset Production environmentThe same agent, pointed at the real systems. Writes wait for a named person in the function that owns the outcome.Where people use the agentThe Mindset HubColleagues, in a browserYour own productAn embedded elementBack in ClaudeOver MCP, where they already work
Build where you build today. The working result comes into a Mindset Test environment, then into Mindset Production once it holds up.
**Mindset gives you environments.** Named partitions inside your Mindset organisation, normally Test and Production, each with its own connections pointed at its own systems. An agent built in Test cannot reach a real system, because the credentials for the real system are not in that environment. - **Nothing is copied for you.** Promoting means making the same change again in Production once it holds up in Test. - **You cannot invite somebody into Test only.** Membership of a Mindset organisation gives a person every environment in it. ## Section 05 — How something gets added Four steps to bring an automation in, done once. After that the agent runs, and every run is recorded.
From an existing automation to a running agent
Once, per automation1Bring it inA Claude or ChatGPT skill, or anMCP server, pasted into Mindset.It proposes a plan and waits foryou to approve it.2Connect the systemCredentials held by Mindset.Operations discovered from thesystem, then enabled one at atime.3Script and testStages in order, each withsomething it must achieve. Everytest criterion has to pass to golive.4PublishTo colleagues in the Mindset Hub,into your own product, or backinto Claude over MCP.Every time it runsA run startsA person, a schedule, your ownsoftware, another Mindset agent, orClaude over MCP.It readsThe operation runs. The answercomes back straight away.It writesNothing happens yet. A person opensa link and approves it.It is recordedWhat started it, every operation itcalled, what came back, how itended.An agent can never approve a write. Not its own, and not another agent's.
The top row happens once per automation. The bottom row happens every time the agent runs.
| Term | What it means in Mindset | | --- | --- | | **Connection** | A link to one outside system, holding its login details. Those details never reach the agent or the model. | | **Operation** | One named thing an agent may do on a connection. Not *the finance system*, but *get the purchase order matching this invoice number*. | | **Run** | One execution, recorded from whatever started it to however it ended. | A Claude skill becomes a Mindset agent. An MCP server becomes a Mindset connection, with each of its tools as a named Mindset operation. ## Section 06 — The rules we enforce Ten rules. Mindset makes seven of them true on its own, whether or not anybody has read the policy. Three are conventions your team has to hold, and those are the three worth writing down.
The ten rules, and who enforces each
Mindset enforces theseSeven rules that hold whether or not anyone reads the policyEverything reaching an outside system goes through a named operationCredentials sit with Mindset and never reach the agent or the modelA write waits for a person, unless your Mindset organisation deliberately turns that offAn agent can never approve a write, its own or another agent'sEvery action is recorded against the run it belonged toWhat you build in one Mindset environment cannot see anotherNothing goes live until every test criterion passesYou enforce theseThree conventions. These are the ones to write downEach agent has a named owner and a backupPersonal tools stay personal and are not registeredAgents nobody uses get archived
A rule Mindset enforces needs no policing. A convention does.
The rule to decide deliberately is rule three. Every Mindset operation is marked read or write. - **A read** fetches information. The agent calls it and the call happens. - **A write** changes something. The agent calls it and nothing happens yet. It records exactly what it wants to send, the run carries on and finishes, and the change waits behind a link a person opens. Write approval is on by default. It can be turned off, but only for the whole Mindset organisation, never per agent and never per operation. When that setting stands in for a person, the record names the setting rather than inventing an approver. Leave it on for the first month. ## Section 07 — Getting started Two phases. Find what your people already built, then get one thing live with us alongside. **Finding them.** Three routes, and most companies need all three. - Ask your AI leads and power users. They tend to have a list already. - Run one session per function. The question is not what could be automated, but what somebody has already automated and is quietly relying on. - Read what your tenant admin tooling and data loss prevention already found. Most is noise. The entries touching a real system are the ones you want. **Qualifying them.** Three questions per item, twenty minutes for the whole list. 1. Can it reach what it needs through an API, a database, a spreadsheet or an MCP server? 2. Is it acceptable for a person to approve the changes it makes? 3. Can you describe the work as stages, each with something it must have achieved? Three yeses and it moves as it is. Start with the biggest, because you already know how it should behave. **The first sprint** runs about three weeks with us, then a month of watching. ## Section 08 — Running it week to week Most of the watching is done for you. Mindset holds every write until a person approves it, blocks any version that fails a single test criterion, marks every agent and connection as working, failing, idle or never used, and compares what each agent was granted against what it actually called.
What Mindset holds on its own
Mindset does this on its ownNothing goes live that fails atestEvery agent carries behaviour criteria written when it was built. A new version goes live only if every singleone passes. Not most. Every one.Every resource carries its ownstateWorking, failing, idle or never used. Idle and never used are separate, because an agent that has never run hasa trigger problem and one that stopped has a different problem.Granted is compared with calledWhat each agent was given, against what it actually reached for. Something granted and never called means astage is being skipped. Something called and never granted is worth looking at the same day.Nothing reaches a real systemunaskedA write is held until a person opens the link and approves it. An agent can never approve one.A person does thisTen minutes each morningRead what failed, what is waiting on an approver, and what ran but changed nothing.Thirty minutes each weekRead the new agents and new operations from the week, and the granted against called comparison.One named person, and a backup. That is the whole standing commitment.
What runs without anybody asking, and the two habits a person keeps.
A person does the rest: ten minutes each morning reading what failed, what is waiting on an approver, and what ran but changed nothing; thirty minutes each week reading the new agents and operations, and the granted-against-called comparison. One named person, and a backup — that is the whole standing commitment. The guide to the record is itself an agent, published to your team over MCP, so the question gets asked from Teams, Copilot, Claude or ChatGPT and answered there — "did the invoice agent run this morning?", "what failed yesterday, and at which step?", "what is waiting on an approver?", "which agents can write to finance?", "what did we spend, and on which agent?" — as a sentence, a number, a table, or over OpenTelemetry into your own tooling. > **Name one owner.** In the same Cloud Security Alliance survey, responsibility for agent security was split across security, engineering and IT, IAM teams were rarely the primary owner, and some organisations reported no identifiable owner at all. On who is accountable when an agent causes an incident: 28 per cent said security or IT, 25 per cent engineering, 18 per cent the business owner, and 15 per cent did not know. --- ## What Mindset does not do **Section:** Start here ## Mindset does not find agents you do not know about If somebody builds an agent in Claude, ChatGPT or Copilot Studio and connects it straight to a system, Mindset never sees it. Discovery belongs to your tenant admin tooling, your data loss prevention and your threat detection platform. Discovery and Mindset work together. Discovery tells you something exists. Mindset is the record of what was allowed. Anything discovered and absent from Mindset reached its system by some other route, and that absence is the finding. ## Mindset does not drive a browser There is no browser automation, deliberately. An automation that works by clicking through a screen leaves no record anyone can check afterwards. If the only way to reach a system is through its interface, that automation stays where it is until the system has an API, a database, a sheet or an MCP server. ## A run does not resume If a run stops part way, it stops. Nothing picks it up from where it got to. Re-running is a decision a person makes, and a re-run is a new run that starts from the beginning. Anything that has to survive a failure and continue belongs in a queue, not here. ## Mindset does not replace your identity or access tooling Who may use which model, who is licensed, who is in which group. That stays where it is today. Mindset governs what a sanctioned agent may do once it exists. ## Mindset does not fix an automation nobody had defined If the work has never been written down as steps, Mindset will not discover them for you. The three-question test for whether something is ready to move is in [Setting up Mindset](/docs/guides/setting-up-mindset). --- ## Where we fit with what you own **Section:** Getting oriented
The four layers of agent governance
SurfaceWhere people work. The assistants you already licence.You already own thisExecutionWhat a sanctioned agent may do. Named operations, its own identity, recorded, tested.MindsetDiscoveryWhat exists that nobody sanctioned. Tenant admin tooling, data loss prevention, threat detection.You already own thisAccessWho may use which model. Identity provider, licensing, vendor admin consoles.You already own this
You already own three of these four layers. The fourth is what this guide is about.
## First, the words Half the confusion in a mixed estate comes from four words that exist in every product and mean something different in each. Worth settling before anything else. | Word | In Microsoft | In Claude and ChatGPT | In Mindset | | --- | --- | --- | --- | | **Environment** | A Power Platform environment. Holds apps, flows and data. Membership set per environment. | No equivalent. One organisation or workspace, and nowhere to build that is not live. | A named partition, normally Test and Production, with its own connections. Unrelated to Power Platform. | | **Organisation** | Your Entra tenant. | Your Claude organisation or ChatGPT workspace, where connectors are allowed or not. | Your Mindset organisation. The only boundary, and where the write-approval setting lives. | | **Connection** | A Power Platform connection to a connector. | A connector, allowed for everyone or nobody. | A link to one outside system, holding its login details. Six kinds, including an MCP server and a model. | | **Operation** | A connector action. | A tool on a connector, or a skill's step. | One specific named thing an agent may do on one connection, marked read or write. | > From here on, every use of these words is prefixed with the product it belongs to. Where this guide says environment without a prefix, it means a Mindset environment. ## Access: Who may use which model Microsoft Entra ID for identity and conditional access. The Microsoft 365 admin centre for Copilot seats, the Claude Enterprise admin console, the ChatGPT Enterprise admin dashboard. Training and attestation on top. In most estates this part is finished. Seats are bought, managed and attested. ## Discovery: What exists that nobody sanctioned Tenant admin tooling shows what has been created. Microsoft Purview captures prompts and responses and classifies data. A threat detection platform ingests provider telemetry. A discovery audit in a company of a few hundred people typically turns up several hundred agents. - Some shipped with an ERP deployment and were never asked for. - Most were built by business users somewhere nobody was watching, often the Power Platform default environment. - A handful of MCP server hookups nobody knew about. That is a discovery layer working correctly, telling you something you cannot act on. ## Execution: What a sanctioned agent may do Which systems an agent reaches, which named operations on those systems, whose credentials it uses, what is recorded, and whether a change happens at all without a person seeing it first. Execution is the layer nothing in most estates covers. The agents touching real systems do so under the credentials of whoever invoked them, and some of those people hold administrator rights in production. Nothing in the estate can tell an agent action apart from a human one. ## Surface: Where people work - **The Mindset Hub** for colleagues, in a browser. - **An embedded element** inside your own product, for customers. - **MCP**, so an agent is reachable from somebody's own Claude or Claude Code. All three carry the same grant. An agent reached from Claude has exactly the operations it was given, no more and no fewer than from the Mindset Hub. Environments are the part of this that most often gets confused with something else, so they have an article of their own: [Environments](/docs/guides/mindset-environments). --- ## Environments **Section:** Getting oriented ## What an environment holds Not a copy of your work, and not a permission level. An environment is a wall. Everything built inside one belongs to it and cannot see anything in another.
What a Mindset environment holds
Your Mindset organisationOne organisation. One region. Members belong to all of it.Mindset Test environmentWhere people buildFinance connectionYour test finance systemCRM connectionYour test CRMModel connectionA cheaper model, cappedAgents here cannot see anything in Production, and cannot reach areal system.Mindset Production environmentWhere the work happensFinance connectionYour real finance systemCRM connectionYour real CRMModel connectionThe model you choseWrites here wait for a named person before anything is sent.Nothing crosses on its ownPromoting means making the same change again in Production once it holds up in Test. Membership of the Mindset organisation is access to everyenvironment in it, so there is no way to invite somebody into Test only.Mindset environments are not Power Platform environments. They are separate objects, and one has no effect on the other.
Two environments inside one Mindset organisation. The connections are what make them different, because each points at different systems.
The connections are the important part. A finance connection in your Mindset Test environment holds the credentials for your test finance system. An agent built in Test cannot read a real invoice, cannot raise a real query, and cannot be made to by editing its instructions, because the credentials for the real system are not in that environment. ## What your teams have today Only one of the products your people build in gives them somewhere to build that is not live, and it is not one of the two most of them use.
Is there somewhere to build that is not live?
Microsoft Copilot StudioYesPower Platform environmentsYou probably already have several.Claude EnterpriseNoOne Claude organisationSkills and connectors are on foreveryone or nobody.ChatGPT EnterpriseNoOne ChatGPT workspaceSame shape as Claude.MindsetYesMindset environmentsNormally Test and Production, eachwith its own connections.
For the two that answer no, a Mindset Test environment is the first non-production place those teams have had.
## How the work moves through them Nobody moves tool. People keep building where they build today, and the working result comes into a Mindset Test environment first.
How the work moves through Mindset environments
Where people build todayClaude EnterpriseA skill file, or an MCP server on a laptopMicrosoft Copilot StudioAn agent built inside one of your environmentsChatGPT EnterpriseA custom GPT, a skill, or an actionMindset Test environmentMindset connections pointed at your non-production systems. The agent cannot read a real invoice even if its instructions tell it to. Not a Power Platformenvironment, and unrelated to one.You make the same change again. Nothing crosses on its own.Mindset Production environmentThe same agent, pointed at the real systems. Writes wait for a named person in the function that owns the outcome.Where people use the agentThe Mindset HubColleagues, in a browserYour own productAn embedded elementBack in ClaudeOver MCP, where they already work
Build where you build today. The result comes into a Mindset Test environment, then into Mindset Production once it holds up.
The agent is published back out to wherever people already work: the Mindset Hub for colleagues, an embedded element in your own product, or back into Claude over MCP. All three carry the same grant. --- ## What we bring in **Section:** Getting oriented
What comes in, and what it becomes
What your people already builtWhat it becomes in MindsetA Claude or ChatGPT skillAn instruction file somebody tuned over monthsA Mindset agent your organisation owns, rather thansomething living in one person's account.An MCP serverOften on a laptop, often pointed at something realA Mindset connection, with each of its tools as a namedMindset operation.An API, database or spreadsheetHTTP, Postgres, Google SheetsA Mindset connection. Operations are discovered, andenabling one makes a real call to prove it works.Documents an agent readsA Copilot knowledge source, or a Claude projectA Mindset knowledge connection, so reaching it is recordedlike anything else.
Nothing is built from scratch. Four kinds of starting material, and what each becomes.
## Skills and prompt files A Claude skill, a long instruction file, a prompt somebody has tuned over months. This is the most common starting point and the easiest. The file gets pasted into Mindset, which reads the connections and operations that already exist in your Mindset organisation, proposes what to build, and waits until somebody approves the plan. What comes out is a Mindset agent owned by your Mindset organisation, rather than something living in one person's Claude account. The instructions become the agent's standing instructions, the stages become an ordered list the agent works through, and anything the agent needs to reach becomes a connection with named operations on it. ## MCP servers Someone has an MCP server running, often on a laptop, often pointed at something real. The shape is the same. The server becomes a connection and each of its tools becomes an operation on that connection. One thing a person does rather than Mindset. An MCP server publishes its tools without reliably saying which of them change the world, so somebody marks each one read or write. Err towards write. - A read marked as a write costs somebody an approval click. - A write marked as a read means an agent changed something in one of your systems and nobody was asked. ## APIs, databases and spreadsheets An HTTP API, a Postgres database, a Google Sheet. - Operations are discovered rather than hand-written. - Enabling one makes a real call to prove it works, and the output shape comes from that real response rather than from the documentation. - For a database, Mindset checks what the credentials you supplied are allowed to do before enabling anything, and proves the query works without changing any data. ## Knowledge bases Documents a Mindset agent needs to reason over. In Mindset a knowledge base is a connection like any other, so an agent reaching it does so through something named and recorded. Material that is currently a Copilot Studio knowledge source or a Claude project becomes a Mindset knowledge connection. > **What does not come across:** anything that works by driving a browser, and anything that has to resume after a failure. Both are boundaries rather than roadmap items. --- ## The policy you write **Section:** Policy This is a policy **you** write and own. It is not a Mindset setting and not something we apply on your behalf. What follows is the version we see working across the companies Mindset runs in: ten rules, phrased so you can lift them into your own document. The reason it is short is that seven of the ten are already true. Mindset enforces them whether or not anybody has read the policy, so writing them down is documenting a control rather than issuing an instruction. The remaining three are conventions your team has to hold, and those are the ones that decay if nobody is watching.
The ten rules, and who enforces each
Mindset enforces theseSeven rules that hold whether or not anyone reads the policyEverything reaching an outside system goes through a named operationCredentials sit with Mindset and never reach the agent or the modelA write waits for a person, unless your Mindset organisation deliberately turns that offAn agent can never approve a write, its own or another agent'sEvery action is recorded against the run it belonged toWhat you build in one Mindset environment cannot see anotherNothing goes live until every test criterion passesYou enforce theseThree conventions. These are the ones to write downEach agent has a named owner and a backupPersonal tools stay personal and are not registeredAgents nobody uses get archived
The ten rules. Seven hold on their own. Three are yours to keep.
## Rule one: Everything reaching an outside system goes through a named operation Not a system, an operation. The unit of permission is *get the purchase order matching this invoice number*, not *the finance system*. Permission at operation level is what makes the rest of the policy possible, because a list of operations is something a person can actually read and sign off. ## Rule two: Credentials sit with Mindset's secure secret keeper, never with the agent or in a chat The login details for a connection are held by Mindset. They do not reach the agent and they do not reach the model. An agent running under a named person's credentials is the single most expensive thing to unwind later, because every log you own will attribute its actions to that person. ## Rule three: A write waits for a person The agent records exactly what it wants to send, the run finishes, and the change waits behind an approval link. This is on by default. It can be turned off, but only as a deliberate setting on the whole Mindset organisation, and the record then names the policy rather than inventing an approver. Two design rules make approval worth having. Make one change rather than fifty, because one approval covering a batch is a decision somebody actually reads and fifty separate ones get waved through. And put enough in the approval that whoever reads it does not have to go and look anything up. > **Worth knowing before you rely on it.** A study of more than 40,000 agent runs and roughly 409,000 approve-or-deny decisions found around one in three dangerous commands were approved, and that people catch what looks dangerous rather than what is. The EU AI Act names automation bias in Article 14, requiring that anyone overseeing a system stays aware of it. Showing what is about to happen, rather than asking yes or no, is materially better than a bare prompt. Volume is still the thing that breaks it. ## Rule four: An agent can never approve a write Not its own, and not another agent's. There is no setting on an operation that changes the rule. The rule matters more as agents start calling other agents, because that is the point at which a chain of approvals could otherwise close on itself. ## Rule five: Every action is recorded against the run it belonged to Agents cannot skip the record. Every run can be read back afterwards step by step, exactly as it happened. If your evidence has to live somewhere else, spans export over OpenTelemetry into whatever your team already audits. ## Rule six: What is built in one Mindset environment cannot see another Mindset Test and Mindset Production are separate partitions with separate Mindset connections. Nothing crosses on its own. Separation is what lets people build against real shapes without touching real systems. ## Rule seven: Nothing goes live until every test criterion passes Not most of them. Every one. A better average never overrides a single failure. The criteria are written by whoever built the agent, so the review that matters is of the criteria themselves, and that goes to the system owner once before the first version goes live. The set that passed is recorded against the version. ## The three rules that are yours These are conventions rather than controls, and they are the ones that decay. - **Each agent has a named owner and a backup.** There is no owner field, so the owner and backup go in the agent's description, in the same format every time. The failure is never that nobody can do the job. It is that everybody could, so nobody does. - **Personal tools stay personal.** Do not register them, do not review them. A policy that governs the agent tidying somebody's meeting notes will not be taken seriously on the agent touching bank details. - **Agents nobody uses get archived.** Anything that has not run in thirty days is a question, not a fixture. > Write the policy as ten lines, not ten pages. Seven of them are already true. Three of them are promises your team is making, and being honest about which is which is what stops the document being decoration. ## Deciding what the policy covers The rules above are what a governed agent lives under. This is how you decide which agents those are, and it is the half that stops a policy being ignored.
Deciding what belongs in Mindset
For each automation your people have already built, ask:1Does the automation change anything?Reads are recoverable. Writes are not.2Can that change be undone?Updating a status and issuing a payment are both writes.3Does anyone but the builder rely on the output?The second person turns it into infrastructure.4Does it run under a person's credentials?This is the one that removes your ability to investigate.5Does it run when nobody is watching?A scheduled agent has nobody to notice it going wrong.6Does it touch data you would have to report on?Personal or regulated data, or anything with a notification duty.Any single yesThe automation belongs in Mindset. Reads run.Writes wait for a person.Six times noSomebody's personal tool. Leave it alone. Donot register it, review it or ask anyone todeclare it.
Six questions per automation. Any single yes and it belongs in Mindset.
Six times no and the automation is somebody's personal tool. Leave it alone. Do not register it, do not review it, do not ask anyone to declare it. A policy that governs the agent tidying somebody's meeting notes will not be taken seriously on the agent touching bank details. ## The triggers that appear later An agent's risk profile changes during its life, not at the start of it. A policy that only assesses at the start will be wrong within months. - **Crossing a boundary.** An agent that starts reaching between domains, or between systems that were previously separate, has changed shape even if nothing about its instructions changed. - **Delegation.** An agent that calls other agents is a different risk from one that does not. Mindset records a delegated run inside the run that called it, so delegation is visible rather than inferred. - **Compositional risk.** Two operations that are each safe and are not safe together. Nobody catches this by reviewing operations individually. - **Drift.** The agent behaves differently without anyone changing it. The comparison of what an agent was granted against what it actually called is where this shows up first. - **Failure history.** Something that has failed before deserves tighter handling than something that has not. - **Volume and velocity.** A change in how often the agent runs, or how much it does per run, against its own baseline. - **Facing outward.** Anything a customer or a third party sees. ## Three tiers, if you need something more graduated In or out is enough for most organisations. If you need more, three tiers is the smallest version that still means something.
Three tiers of agent
TierWhat it may doWho signs it offReviewedRead onlyFetches and reports. Changes nothing.Whoever builds itWhen it changesWrites, one at a timeEvery change waits for a named person to approve it.The business owner and thesystem ownerQuarterlyWrites on a standing policyA standing setting on your Mindset organisation standsin for the person. The record names that setting.Security and the systemowner, in writingMonthlyMost agents stay in the middle tier for good. Moving to the third is a decision somebody signs, not a setting somebody flips.
Sign-off and review frequency both scale with the tier.
This is the Cloud Security Alliance's autonomy scale cut down to the levels you can operate. Their own position is that the top of their scale is not appropriate for enterprise deployment today. Everything Mindset does by default is the middle tier. --- ## Setting up Mindset **Section:** Implementing it ## Phase one: find what already exists The material is already there, in skill files, prompt documents, MCP servers on laptops and spreadsheets somebody maintains by hand every Friday. You are looking for what is already working, not for ideas nobody has tried. - **Ask the people who already build.** Whoever your AI leads or power users are. They usually have a list already. - **Run one session per function.** Finance, operations, HR, an hour each. The question is not what could be automated. It is what somebody has already automated and is quietly relying on. - **Read what your discovery tooling found.** Tenant admin tooling and data loss prevention already have a list. Most of it is noise. The few entries touching a real system are the ones you want. Write them down in one list with the person who built each one. That list is the asset, and it usually surprises people. ## Then qualify the list Three questions per item, twenty minutes for the whole list. 1. Can the automation reach what it needs through an API, a database, a spreadsheet or an MCP server? 2. Is it acceptable for a person to approve the changes the automation makes? 3. Can you describe the work as stages, each with something it must have achieved? Three yeses and the automation moves as it is. Start with the biggest, because you already know how it should behave, which makes it the cheapest thing to learn on. Anything driven by clicking through a screen, and anything that has to pick itself up after a failure, does not move. ## Phase two: the first sprint One automation, taken all the way through, with us alongside. Not for the automation. For the fact that your team has done every step once and does not need us for the second one. | Phase | What happens | Who | | --- | --- | --- | | **Setup** | Create the Mindset organisation, its environments and its governance settings. Region, retention, whether OpenTelemetry export is on, whether write approval stays on. | Your Mindset owner, with us | | **Connect the first system** | Credentials, discover operations, enable the ones the automation needs, mark each read or write. The conversation with whoever owns that system. | System owner, with us | | **Build and test** | The automation comes into Mindset, which reads what already exists in your Mindset organisation, proposes a plan and waits. Then the stages, then the behaviour criteria. Also where you settle which model the agent uses, by running the same criteria against a cheaper model and taking the cheapest that still passes everything. | The person who built it, with us | | **Live and watched** | Make the same change in your Mindset Production environment and publish it. Then it runs, and you read the runs daily. Nothing else gets built until you can say what a normal run looks like. | Mindset owner and system owner | ## What the first one costs you The effort scales with what you are bringing across rather than with Mindset, so treat these as the shape rather than a quote. | | Roughly | What moves it | | --- | --- | --- | | **Governance decisions** | About an hour | Fixed. Same for everyone. | | **Each system you connect** | One conversation with the owner, then setup | How many systems the automation touches, and how many operations it needs on each. An API with good documentation is quick. A database where nobody knows the schema is not. | | **The automation itself** | Days, not weeks, if it already works | How many stages, whether the work is already written down, and whether it writes or only reads. | | **Your team's time across the sprint** | A couple of days for the Mindset owner, plus the builder | Mostly whether the people who own the systems are available. | The second automation costs a fraction of the first, because the connections and the operations already exist. That is the number worth planning around, not the first one. ## Decisions to make before day one - **Region.** Fixed when the Mindset organisation is created and not changeable afterwards. - **Mindset environments.** At least Test and Production. Names cannot be changed later. - **Write approval.** Leave it on for the first month. Revisit once you can describe what a normal run looks like. - **Retention and export.** How long records are kept, and whether the full record goes to your own logging tooling over OpenTelemetry. - **Personal agents in Mindset.** Off by default. Turning them on gives people somewhere in Mindset to experiment that holds no connections and reaches nothing. This is where personal productivity agents belong. > The thing to protect in the first month is the second automation, not the first. If your team needs us for it, the sprint did not work, however well the first one runs. --- ## Running Mindset without spending hours on it **Section:** Running it ## What runs without anybody asking
What Mindset holds on its own
Mindset does this on its ownNothing goes live that fails atestEvery agent carries behaviour criteria written when it was built. A new version goes live only if every singleone passes. Not most. Every one.Every resource carries its ownstateWorking, failing, idle or never used. Idle and never used are separate, because an agent that has never run hasa trigger problem and one that stopped has a different problem.Granted is compared with calledWhat each agent was given, against what it actually reached for. Something granted and never called means astage is being skipped. Something called and never granted is worth looking at the same day.Nothing reaches a real systemunaskedA write is held until a person opens the link and approves it. An agent can never approve one.A person does thisTen minutes each morningRead what failed, what is waiting on an approver, and what ran but changed nothing.Thirty minutes each weekRead the new agents and new operations from the week, and the granted against called comparison.One named person, and a backup. That is the whole standing commitment.
Four things Mindset holds on its own, and the two habits a person keeps.
The first of the four is the one people do not expect. - Every Mindset agent carries behaviour criteria, written when it was built, describing what it must do and must not do. - A new version goes live only if every single criterion passes. Not most of them. - The set that passed is recorded against that version, so months later you can see what a given version was held to. The criteria are written by whoever built the agent, so the review that matters is of the criteria themselves. That goes to the system owner, once, before the first version goes live. What none of it does is notice on your behalf that an agent which ran every morning for six weeks stopped on Tuesday. That is what the two habits above are for. ## Asking without signing in Neither habit means learning new software. The guide to the Mindset record is itself an agent, published to your team over MCP, so the question is asked wherever they already work and the answer comes back in the same window.
Your leads can manage the Mindset app without logging in
You ask from where you already workMindset answers over MCPYou get backMicrosoft Teams, Copilot, Claude orChatGPTSecurity, IT, whoever buildsDid the invoice agent run this morning?What failed yesterday, and at whichstep?What is waiting on an approver?Which agents can write to finance?What did we spend, and on which agent?The Mindset guideAn agent published to your team. It readsand changes nothing.RunsEvery run, what started it, everyoperation it called, how it endedAuditWhich agent could do what, under whoseidentity, and who approved each changeHealthWorking, failing, idle or never usedCostSpend by agent, model, vendor and personAn answer, there and thenIn the same window they asked inIn wordsA sentence that answers the question,with the run or agent namedAs a numberA figure you can put in a board pack or abudget lineAs a tableAsk for it laid out, then paste it into areportInto your own toolingThe full record sent over OpenTelemetryto whatever you already audit
Ask on the left, answer on the right. Nobody signs in to Mindset to find out what happened.
The same route produces the evidence. - Ask for a month of approvals as a table, and paste it into a board pack. - Ask which agents hold write operations and on what, and hand the answer to an auditor. - Or send the full record over OpenTelemetry into whatever your team already retains, and never ask at all. ## What intelligence Mindset offers for different security, ops & leadership teams Security, operations and leadership look at the same record and arrive with different questions. Every one is answered from what was recorded, and every one can be asked from Teams, Copilot, Claude or ChatGPT.
What each team looks at
SecurityCould it have done something it shouldnot have?Which agents hold write operations,and on whatWhat each agent was granted againstwhat it actually calledEvery action, against the run itbelonged toWhat was approved, by whom, and howquicklyOperationsIs it working?Calls, failures and average time, peragent and per connectionWhich agents are working, failing,idle or never usedWhere a run stopped, and what it iswaiting forEvery run replayed in order, exactlyas it happenedLeadershipIs it worth it?Total spend, and cost per call aftereach changeSpend broken down by agent, model,provider and personHow many agents are live and how manypeople use themTime returned against the baseline youcaptured first
The same record, read three ways.
### If you are responsible for security Your question is what an agent could have done, and what it did. | Question | How often | | --- | --- | | Which agents hold write operations, and on what? | Monthly, and after any change | | What was each agent granted, and what did it actually call? | Weekly | | Is anything running under a person's credentials? | Monthly | | What was approved, by whom, and how quickly? | Weekly | | Show me every action in this run, in order | During an investigation | | Which runs did this run hand off to other agents? | During an investigation | | Get the raw traces into our own tooling | Set up once | The one people do not think to ask is the second: what each agent was granted, against what it actually called. - **Granted and never called** means a stage is being skipped in practice. - **Called and never granted** is worth looking at the same day. It usually means a run delegated to another agent that came with resources of its own. ### If you run the operation Your question is whether the agent is working. | Question | How often | | --- | --- | | Did the agent run at all? | Daily | | Did the agent run as often as it should have? | Daily | | How far did the agent get? | Daily | | What is the agent waiting for? | Daily | | Calls, failures, successes, average time, last used | Weekly | | What is working, failing, idle or never used? | Weekly | | Which version ran? | When something changed | Idle and never used are separate states on purpose. Something that has never run has a trigger problem. Something that used to run and stopped has a different problem. ### If you are paying for Mindset Your question is whether the spend was worth it. | Question | How often | | --- | --- | | Total spend, and how it moved | Monthly | | Average cost per call | Monthly | | What does each agent cost? | Monthly | | What are we paying for capability we may not need? | Quarterly | | How is spend split across suppliers? | Quarterly, before a negotiation | | Who is using it? | Monthly | | How much of the spend is cached, and what would the same work cost without caching? | Monthly | | How many agents are live, and how many people use each one? | Quarterly | Two numbers Mindset cannot give you. Time returned against the manual baseline, and how many people can now do something only one person could do before. Both come from your own systems, and the baseline has to be captured before the agent goes live because nobody can reconstruct it afterwards. > **On caps and exports.** A cost cap is set on the model connection, alongside the model, a token ceiling, PII handling and data residency, rather than on each agent. Set a cap on anything scheduled before leaving that agent alone for a month. There is no CSV download of runs. If you want the data in your own tooling, that is the OpenTelemetry export. ## The four jobs, and how often each comes up Permissions attach to agents rather than people, so this is not an access matrix.
The four jobs in running Mindset
Build itWhoever knows the process beingautomatedOnce per automationDecide which operationsexistWhoever is accountable for thesystem being touchedOnce per systemApprove writesA named person in the functionthat owns the outcomeAs they arriveWatch itOne named Mindset owner, and abackupTen minutes a day
Only the last of the four is a standing commitment.
**Build it.** Anyone. They know the process being automated, which is the part that cannot be delegated. **Decide which operations exist.** Whoever is accountable for the system being touched. A **Mindset operation** is one specific named thing an agent may do on one system: *get the purchase order matching this invoice number*, rather than access to the finance system. The system owner is asked once which of those should exist, and that list is the boundary for every agent built against that system afterwards. **Approve writes.** A named person in the function that owns the outcome. Finance approves finance writes, not the AI team. A Mindset owner approving writes across every system becomes a bottleneck and a rubber stamp at the same time. **Watch it.** One named Mindset owner and a backup. The only standing commitment. ## The cadence | When | Who | What | | --- | --- | --- | | **Daily, ten minutes** | Mindset owner | What failed, what is waiting on an approver, what ran and changed nothing. | | **Weekly, thirty minutes** | Mindset owner and backup | New agents and new operations from the week. What each agent was granted against what it actually called. | | **Monthly, an hour** | Add security and a system owner | Cost by agent and model. Agents that have not run in thirty days. Anything your discovery tooling found with no sanctioned route. | | **Quarterly** | Add leadership | Time returned against baselines. Owners and backups reconfirmed. Agents archived. | > Put the monthly hour in the diary before the first agent goes live. The failure mode is not that the review is hard. It is that nobody schedules the review, and the first one happens after an incident. --- ## Your plan on a page **Section:** Reference Answer the bullets below and you have the page. ## 1. What are we protecting against? - What is the specific loss, in one sentence, not the category? - What would it cost us if it happened once? - How would we find out that it had happened? ## 2. What is true today? - How many agents exist, and where did they come from? - How many of them can write to a real system? - How many run under a person's credentials rather than their own identity? - Who could tell us, if we asked right now? ## 3. What are the rules? - What ten rules will a governed agent live under? (See [The policy you write](/docs/guides/the-policy-you-write).) - Which of them are enforced by the platform and need no policing? - Which are conventions our team is promising to hold? - Who signs off that the list is complete? ## 4. What is in scope, and what is deliberately not? - Which of the six questions puts something in scope for us? (See [The policy you write](/docs/guides/the-policy-you-write).) - What are we explicitly leaving alone, and have we said so in writing? - Where do personal productivity agents live instead? ## 5. Who does what? - Who is the named owner, and who is the backup? - Who approves writes for each system, and are they in the function that owns the outcome? - Who decides which operations exist on each system? - Who is accountable when something goes wrong? ## 6. How do we get started? - How will we find what people have already built? - Which automation are we starting with, and why that one? - What are the dates, and who is in the room? (See [Setting up Mindset](/docs/guides/setting-up-mindset).) - What does the second one cost, once the first is done? ## 7. What will we measure? - Which numbers come from Mindset? (See [Running Mindset without spending hours on it](/docs/guides/running-mindset-day-to-day).) - Which number comes from our own systems, and have we captured the baseline before anything goes live? - What does the monthly review look at, and who is in it? - What would tell us early that this is not working? ---