Playbook

Stop Building AI Assistants. Build an AI Management Team With Grok Bot.

Build enough useful bots and you become the busiest person in the system. Each bot is good at its job. You are the one deciding which bot got which job, pasting context between them, and checking who did what.

By Mitchell Keller18 min readGrok Bot

Marble plinths arranged as a hierarchy on a lakeside terrace, with a cobalt line running from a ship's helm through one domain to three instruments
On this page

The fix wasn't a better bot. It was a management layer. Grok Bot makes that layer easy to build.

By the end you'll have a structure you can copy: one bot you talk to, one bot that owns GTM, a bench of experts, recurring jobs, and clear rules for what the system may do without asking.

You
First Mateyour only front door
GTM Second Mateowns outbound and pipeline
Content Second Matenewsletter, social, site
Ops Second Matebilling, tooling, admin
Consultantsjudgment on demand
Specialistsnarrow jobs
Tools and other AIconnectors, browser, terminal
Routinesjobs that run without a prompt
Always-on boxtools off your laptop
Work surfacesCRM, inbox, campaign platform
Approval gates between every layer and the outside world
The whole model. The GTM branch is the one this article builds.

The map

Here is the whole model before any detail:

  • You talk to one bot: the First Mate.
  • The First Mate hands work to Second Mates. Each one owns a domain, like GTM.
  • Second Mates pull in consultants for judgment and tools for execution.
  • Routines give each owner jobs that run without a prompt.
  • Work surfaces, like your CRM and campaign platform, let the system see real state.

We start with the part you can set up this afternoon and add one layer per section. Each layer fixes the bottleneck the previous one exposes.

Layer 1: The First Mate is your only front door

The simplest change has the biggest payoff. Pick one bot and make it the only one you talk to by default. Its job is not to know everything. Its job is to know who does what and send work there.

Without it, you are the router. You remember which bot handles research, which one writes, and which one has the campaign context. That routing is the context switching you were trying to hand off.

Setup: create the First Mate

In Grok Bot, you create a Bot by giving it a short name, one primary job, and a description of how it works (docs). For the First Mate:

  1. Name: First Mate.
  2. Primary job: take every request from me, decide who owns it, delegate, and report back.
  3. Paste in the operating template below and edit it to your business.
  4. Test it with one messy request (example further down) before you add any other bots.
first-mate.txt
1Who I am

ROLE and ABOUT ME: the job, the business, the quarter's goals, how you like updates.

2Who does what

One line per Second Mate. If nobody owns it, say so and propose an owner.

3Delegation rules

One domain, one owner. Split mixed requests. Anything that sends, spends, publishes, or deletes comes to you.

4Report format

What you asked for, who took each part, what is done, waiting, or needs a decision.

The template below, in four zones.

Example First Mate operating template

Template
ROLE
You are my First Mate. I talk to you, not to the other bots.
Your job is coordination: route work to the right owner, track it,
and report back. You do not do domain work yourself if an owner exists.

ABOUT ME
- Business: [one line on what we sell and to whom]
- This quarter's goals: [2-3 goals with numbers]
- How I like updates: short, decision first, evidence after.

WHO DOES WHAT
- GTM Second Mate: campaigns, lists, messaging, replies, pipeline.
- Content Second Mate: newsletter, social, site copy.
- Ops Second Mate: billing, tooling, admin.
If nobody owns a request, tell me and propose an owner. Do not guess.

DELEGATION RULES
1. One domain, one owner. Never give the same job to two bots.
2. Split a mixed request into parts and send each part to its owner.
3. Pass the owner the goal, the deadline, and what done looks like.
4. Anything that sends, spends, publishes, or deletes comes to me first.

REPORT FORMAT
- What I asked for
- Who took each part
- What is done, what is waiting, what needs my decision

Kun Chen's open-source firstmate is a useful companion here. It is an "agent distro" built on the same idea for coding agents: you talk to one agent and it runs the crew, each worker in its own copy of the code. Grok is listed among its supported harnesses. Read its AGENTS.md for templates you can adapt to your First Mate.

The messy request test

Give your First Mate something that crosses three owners:

"The fintech campaign replies dropped this week, the newsletter is due Thursday, and I think we're double paying for an enrichment tool. Sort it out."

A working First Mate does not try to answer all of it. It sends the reply drop to GTM, the newsletter to Content, and the billing question to Ops, then comes back with one status list.

The fintech campaign replies dropped this week, the newsletter is due Thursday, and I think we're double paying for an enrichment tool. Sort it out.You
First Mate routes
  1. “replies dropped”GTM Second Mate
  2. “newsletter Thursday”Content Second Mate
  3. “double paying”Ops Second Mate
One status report, to You
  • GTM: investigating reply drop
  • Content: draft due Thursday
  • Ops: needs your decision
Three owners, one status list back. Report lines are illustrative.

You do not need to remember which bot handles which task. You talk to one management agent.

Layer 2: Second Mates own domains, not questions

Once the First Mate routes work, the next question is where it lands. A Second Mate is a bot that owns a domain: its projects, context, tools, and specialists. GTM is the best one to build first, because GTM work is a chain of steps that all need the same context.

The rule is one domain, one owner. If three bots all do "research and strategy," they overlap, contradict each other, and you end up judging between them. Give each domain one owner and the overlap disappears.

Example GTM Second Mate operating template

Template
ROLE
You are the GTM Second Mate. You own outbound and pipeline for [company].
You take work from the First Mate and report back to it, not to me,
unless a decision needs me.

WHAT YOU OWN
- Who we sell to: [segments, titles, company size, disqualifiers]
- Offer and messaging: [current offer, proof points, banned claims]
- Active campaigns: [name, segment, status, goal]
- Experiments: [what we're testing, success threshold, end date]
- Pipeline: [where deals live, what "stalled" means, e.g. no touch in 10 days]

YOUR TOOLS AND SPECIALISTS
- List building and enrichment: [tool / connector]
- Campaign platform: [tool / connector]
- CRM: [tool / connector]
- Consultants you may call: [see bench below]

OPERATING RULES
1. Start every project from a written brief: goal, segment, volume, deadline.
2. Show your work: counts at each step, and what you dropped and why.
3. Draft anything outbound. Never send, launch, or import without my approval.
4. When a metric moves more than [X]% week over week, investigate before reporting.

DAILY ANSWER TO "WHAT SHOULD WE FOCUS ON TODAY?"
Top 3 items, each with: the evidence, the owner, the next step.

Ask it "what should we focus on today?" If the answer draws on your campaigns, experiments, and pipeline rather than generic advice, the Second Mate is working.

A bot can own a domain, not just answer questions about it.

Layer 3: Run one real GTM project end to end

This is where the hierarchy proves itself. Give the GTM Second Mate one concrete outcome and let it coordinate every step. The leverage is not writing emails faster. It is removing the human handoffs between the steps.

An example brief: "Build a campaign for Series A fintech companies with 20 to 200 employees whose head of sales was hired in the last six months. Target 500 verified contacts. Draft three messaging angles. Have it ready for my review Friday."

  1. Brief1 goal
  2. Researchsegment notes
  3. Accountsexample: 1,200
  4. Contactsexample: 2,100
  5. Qualificationexample: 640 qualified
  6. Messaging3 angles
  7. Campaign draftexample: 500 verified
  8. Your approvalbefore launch
  9. Launchafter yes
  10. Monitoringroutines
One goal and one context moving through every step. All counts are examples, not results.

The Second Mate runs the chain: research on the segment, account sourcing, contact finding, email validation, qualification against your rules, messaging, and a campaign draft. At each step it reports counts and what it dropped. You see one deliverable, not seven tool sessions.

For the list-building steps, the Second Mate can hand work to a specialist bot. Clay Desk by Mitchell is one I built for this, and you can use it. Its page says it runs Clay work end to end: it builds tables, runs enrichments and functions, and uses a real browser when a step only exists in the Clay app. It shows a cost estimate and asks before spending credits, which is the approval boundary this whole article builds toward.

It does not need to be the best tool at every step. Some steps run through connectors, some through other tools. Its value is keeping one goal and one context moving through all of them.

The leverage is coordinating the workflow, not doing any one task slightly faster.

Layer 4: Consultants for judgment calls

Execution problems have answers. Judgment problems have tradeoffs. When the campaign above gets a 1% reply rate, "run the next step" is not the fix. You need people who see the problem differently.

Consultants are bots or saved contexts that represent a framework, a discipline, or a school of thought. Grok Bot lets Bots share context in group chats and pass work to each other (docs), so the Second Mate can open a temporary room with three or four of them.

Example consultant bench for GTM

ConsultantLensThe question it always asks
Offer skepticValue of the offer to a cold readerWhy would a stranger say yes to this in one email?
List puristTargeting and data qualityAre these the people who can buy, or people who can only recommend?
Deliverability hawkInfrastructure and inbox placementAre these emails reaching the inbox at all?
Buyer proxyThe prospect's point of viewWhat would make me delete this in two seconds?

You can also distill a real expert into a consultant. Distill Anyone is a Grok Bot built for this. Its page says it builds a talkable companion from someone's public YouTube transcripts, adds topic depth over time, and refreshes daily.

The useful part is disagreement. Tell each consultant to argue from its lens and attack the others' reasoning. A room where everyone agrees is one model talking to itself four times.

The room: 1% reply rate
  • Offer skepticA stranger has no reason to say yes to this offer in one email.
  • List puristHalf this list can only recommend, not buy. Fix targeting first.
  • Deliverability hawkCheck inbox placement before you touch the copy.
  • Buyer proxyI'd delete this in two seconds. The first line is about you, not me.
GTM Second Mate synthesizesDecision
Likely cause
List: too many recommenders, not buyers
Next test
Re-cut to budget owners, same copy
Owner
GTM Second Mate, report Friday
Illustrative arguments. The value is the disagreement, then one decision.

Then the Second Mate synthesizes. It does not hand you four opinions. It hands you a decision: the likely cause, the next test, and who runs it.

You can assemble expertise around a problem on demand instead of asking one model to role-play everything.

Layer 5: Grok as the front door to your other AI

You probably already pay for other AI: a coding agent, a research tool, another model you like for writing. The usual pattern is switching between chat windows. The better pattern is treating each one as a capability the hierarchy can call.

Each Grok Bot runs on a cloud computer with a browser, filesystem, and terminal, and uses connectors where available and computer use for everything else (docs). That is what makes this possible: if a tool runs in a browser or a terminal, a Bot can be taught to use it.

In the fintech project, say you want a deeper market scan than a quick search. The Second Mate hands that to a research tool, waits for the result, and folds it into the brief. You never open the other app.

The manager doesn't need to out-research or out-write the specialist tools. It needs to know which one to hand the task to. This is a workflow pattern, not a named Grok Bot feature. How it works depends on each tool's login and terms.

Grok becomes the interface to your other intelligence instead of competing with it.

Layer 5.5: Keep context clean with one Code Mode MCP instead of fifty plugins

Once Grok is the front door to your other tools, the obvious move is to plug everything in. CRM, campaign platform, enrichment, inbox, docs. Each one ships as a plugin or MCP server, and each one adds its tool definitions to the model's context before you type a word.

That is where it breaks. Anthropic's engineering team describes agents connected to hundreds or thousands of tools having to process hundreds of thousands of tokens of tool definitions before reading the request (Anthropic). Every tool result also flows back through the context, even when the agent only needs one field from it. The window fills with schemas and raw data, and the answers get worse. That is context rot.

The fix is called Code Mode. Instead of loading every tool, you give the agent one tool that runs code. The agent looks up the method it needs, writes a few lines that call it, and only the filtered result comes back into context.

Cloudflare introduced it: their Code Mode turns MCP tools into a TypeScript API and has the agent write code against it, because models have seen far more real code than tool-call syntax (Cloudflare). Anthropic's version of the same idea cut one workflow from 150,000 tokens to 2,000, a 98.7% reduction (Anthropic).

Before: every plugin loads its schemas
  • crm.search
  • crm.update
  • crm.notes
  • campaigns.list
  • campaigns.stats
  • campaigns.pause
  • leads.enrich
  • leads.verify
  • inbox.search
  • inbox.reply
  • docs.read
  • docs.write
  • calendar.find
  • billing.invoices
  • sheets.append
  • slack.post
  • +34 more
Your task

50 tool definitions in context before the first word of work

After: one codemode tool
  • codemode
  • describe("crm.search")
  • result: 3 fields x 12 rows
Room for the actual work

Methods looked up on demand, results filtered in code

Illustrative. The point is the ratio, not the exact tool names.

How we run it

Our integrations sit behind one gateway that exposes one codemode tool. The agent never sees fifty schemas. When it needs a method, it calls codemode.describe("<connector>.<method>") to read that one signature, then writes a short script that calls it. Anything that writes to the outside world pauses until a human approves it.

Here is what the agent's code looks like. This is illustrative, not a real API:

Illustrative code
// Illustrative only. The agent looks up one method, then calls it.
const spec = await codemode.describe("crm.searchDeals");
const deals = await crm.searchDeals({ stage: "proposal", updatedWithinDays: 14 });
// Filter before anything returns to the model's context.
return deals
  .filter(d => d.amount > 20000)
  .map(d => ({ name: d.name, owner: d.owner, lastTouch: d.lastActivity }));

The model reads three fields per deal instead of the full CRM payload. The raw data never touches the context.

Set it up

  1. Pick a gateway approach - Use Cloudflare's Code Mode, or an MCP gateway that puts many connectors behind one endpoint. Either works if the agent writes code instead of calling fifty tools.
  2. Register your connectors behind it - CRM, campaign platform, enrichment, inbox. Add them to the gateway, not to the agent.
  3. Expose only two or three tools - A search or describe tool for finding methods on demand, and an execute tool that runs the code. Nothing else goes into the agent's context.
  4. Keep the keys in the vault - The gateway pulls API keys from your secrets manager at run time, the same Infisical setup from Layer 7. No key sits in a prompt or in the agent's code.
  5. Put an approval gate on writes - Reads run directly. Anything that sends, spends, updates or deletes pauses until you approve it. This is the same boundary as Layer 10.
  6. Connect Grok to it - The xAI API supports remote MCP servers by URL (docs). Grok Bot's docs only describe connectors installed from its Marketplace (docs), and do not document adding your own MCP server. So treat this step as workflow-specific: reach the gateway through the always-on box in Layer 7, or through the API.

One tool in context instead of fifty. The Bot spends its window on your work, not on reading tool manuals.

Layer 6: Routines turn prompts into responsibilities

Everything so far still waits for you to ask. Routines change that. A routine tells one Bot when to run a workflow, on a schedule or, where supported, after an event, and it usually runs a saved skill (docs).

Think "this is your job every morning," not "run this prompt every morning." The difference is ownership: the Bot owns the outcome and decides what is worth reporting.

The routine I'd set up first isn't a GTM one. It's a bookmark reviewer. It browses alongside you, learns what you engage with, and turns it into drafts for your next posts, then asks what you think so every piece carries your opinion. It is the fastest way to feel a Bot working like you.

Example routine definitions

Start here: the key routine

Bookmark reviewer and builder

Owner
Content Second Mate
Trigger
Daily, plus whenever I bookmark or save a post
Job
Browse what I bookmarked, liked, and engaged with. Learn what I react to and why. Turn the strongest items into post ideas and drafts that use them as inspiration, and ask for my opinion on each so the next post carries my take, not a summary.
Report
Top 3 ideas with the source post, the angle, and one question for me.
Approval
Publishing anything.
Routine

Morning campaign review

Owner
GTM Second Mate
Trigger
Weekdays 7:30 AM
Job
Check every active campaign against last week. Flag any reply rate or bounce rate that moved more than 25%. Investigate before flagging.
Report
Only if something moved. Otherwise one line: "All campaigns normal."
Approval
Pausing or editing a campaign.
Routine

Reply triage

Owner
GTM Second Mate
Trigger
Every 2 hours during business days
Job
Sort new replies into interested, not now, wrong person, unsubscribe. Draft responses for interested and wrong-person replies.
Report
Count by category, plus drafts waiting for review.
Approval
Sending any reply.
Routine

Stalled pipeline check

Owner
GTM Second Mate
Trigger
Mondays 8:00 AM
Job
Find deals with no activity in 10+ days. Suggest a next step for each.
Report
Top 5 stalled deals by value.
Approval
Any CRM stage change.
Routine

Weekly experiment review

Owner
GTM Second Mate, reporting to First Mate
Trigger
Fridays 3:00 PM
Job
For each running experiment, compare to its success threshold. Recommend keep, kill, or extend.
Report
Always.
Approval
Ending an experiment or starting a new one.
The five definitions below, as cards. Every one reports to an owner and names what needs approval.
Template
ROUTINE: Bookmark reviewer and builder
Owner: Content Second Mate
Trigger: Daily, plus whenever I bookmark or save a post
Job: Browse what I bookmarked, liked, and engaged with. Learn what I react to
     and why. Turn the strongest items into post ideas and drafts that use
     them as inspiration, and ask for my opinion on each so the next post
     carries my take, not a summary.
Report: Top 3 ideas with the source post, the angle, and one question for me.
Approval: Publishing anything.

ROUTINE: Morning campaign review
Owner: GTM Second Mate
Trigger: Weekdays 7:30 AM
Job: Check every active campaign against last week. Flag any reply rate
     or bounce rate that moved more than 25%. Investigate before flagging.
Report: Only if something moved. Otherwise one line: "All campaigns normal."
Approval: Pausing or editing a campaign.

ROUTINE: Reply triage
Owner: GTM Second Mate
Trigger: Every 2 hours during business days
Job: Sort new replies into interested, not now, wrong person, unsubscribe.
     Draft responses for interested and wrong-person replies.
Report: Count by category, plus drafts waiting for review.
Approval: Sending any reply.

ROUTINE: Stalled pipeline check
Owner: GTM Second Mate
Trigger: Mondays 8:00 AM
Job: Find deals with no activity in 10+ days. Suggest a next step for each.
Report: Top 5 stalled deals by value.
Approval: Any CRM stage change.

ROUTINE: Weekly experiment review
Owner: GTM Second Mate, reporting to First Mate
Trigger: Fridays 3:00 PM
Job: For each running experiment, compare to its success threshold.
     Recommend keep, kill, or extend.
Report: Always.
Approval: Ending an experiment or starting a new one.

Per current docs, a Bot can own up to 50 routines, and you should test a routine before enabling it, because "a test run performs real work" (docs).

A recurring responsibility is different from a prompt you manually repeat.

Compute, not chat messages

Once routines run, think about usage differently. The one-line "all campaigns normal" message might sit on top of dozens of steps: several Bots, connector calls, browser sessions, another model, and retries.

What you see

Fintech campaign reply rate down 40%. Likely cause: new segment. Recommend pausing variant B.

What ran to produce it
  1. 3 Bots involved
  2. Connector calls: campaign platform and CRM
  3. Browser steps
  4. 1 external research call
  5. 2 consultant turns
  6. 1 retry
One message, many units of work. Example numbers.

So budget for work performed, not messages. Grok Bot usage resets weekly (docs), which makes a weekly budget the natural unit. Start with fewer routines, watch what they consume, then add more. Plan tiers and limits change, so check Plans and billing before relying on any number.

One small output can represent a lot of work, and that is what you are paying for.

Layer 7: An always-on environment for tools outside Grok

First, what you do not need. Grok Bot work runs on its cloud computer, and the docs say closing the app, your laptop, or your phone does not stop a background turn or a routine (docs). You do not need a server to keep native routines alive.

A VPS (a rented computer that stays on in a data center) matters for a different case: tools that live on your own machine. A command-line tool you configured locally, a script, a browser session on your laptop, or software you want the system to reach at 3 AM. Move those to an always-on computer, and they stop depending on whether your laptop is open.

Keep it simple: a VPS is an always-on computer for capabilities that would otherwise live on your laptop. This is a workflow choice. How a Bot reaches that computer depends on your setup, for example a terminal login or a custom MCP server (a small service that exposes a tool to the Bot). It is not a documented default.

Once more than one agent needs the same API keys, keep them in one secrets manager instead of pasting them into each bot's instructions or leaving them in files on the server. We use Infisical: every agent, on the VPS or anywhere else, pulls the keys it needs at run time from the same place. You rotate a key once and every agent picks it up, and no key ever sits in a prompt or a chat log.

External tools do not have to depend on the laptop you happen to be using.

Layer 8: Let Grok see where the work lives

Until now, the system mostly knows what you tell it. That caps its usefulness, because you become the person copying the state of the business into chat.

Connect the places where GTM state actually lives: CRM, inbox, campaign platform, analytics, a database like Supabase, or your project tracker. Grok Bot ships connectors for tools like Gmail, Google Calendar, Google Drive, Outlook, Teams, and Salesforce, and you can add more through the Marketplace. That list comes from third-party guides, so confirm the current one in the app.

Connect only what a workflow needs. The docs say it directly: "Connect only the tools a workflow needs" (docs).

Now the GTM Second Mate can see what you see: unhandled replies, a campaign whose reply rate fell, contacts missing enrichment, deals aging in one stage.

The AI no longer needs you to copy the state of the business into chat.

Layer 9: Let Grok find the work

This is where the direction of work reverses. Once the system can see real state, and routines check it, it can notice problems before you do.

  1. 1

    You notice, you ask AI.

  2. 2

    AI inspects when asked.

  3. 3

    AI inspects on a routine.

  4. 4

    AI notices, investigates, brings it to you.

  5. 5

    AI notices, delegates, prepares the action.

  6. 6

    AI executes permitted actions, updates the system, reports back.

Steps 5 and 6: approval gates required
Asking is step 2. Routines are step 3. Step 4 is where the direction of work reverses.

Asking a bot to check something is step 2. Routines get you to step 3. Step 4 is the jump: the Second Mate spots that the fintech campaign's reply rate halved, checks whether bounces rose, sees that one new segment accounts for the drop, and arrives with the cause and a recommendation.

Before: you pull
  1. You
  2. notice problem
  3. open bot
  4. paste context
  5. ask
  6. read answer
  7. decide
After: the system pushes
  1. Signal in campaign platform
  2. GTM Second Mate notices
  3. investigates
  4. First Mate
  5. You receive one decision card

The AI can notice the problem and come to you.

Layer 10: Controlled execution closes the loop

The last layer lets the hierarchy act, inside limits you set. The full chain looks like this: a signal appears, the Second Mate notices, investigates, pulls in a consultant if the call needs judgment, prepares the next step, takes the actions it is allowed to take, updates the shared surface, and reports up.

Grok Bot has the controls for this. You can set limits in the request itself, and admins can set auto-review rules of "Ask first" or "Allow automatically," where "Ask first" wins when both match (docs). The docs recommend approval for sending, purchasing, deleting, publishing, or changing production systems.

Approval boundaries for a GTM Second Mate

Do it

  • Read CRM, inbox, campaign statsRead-only, no risk
  • Research accounts, find and validate contactsReversible, internal
  • Draft emails, replies, and campaign copyNothing leaves the building

Do it, tell me after

  • Tag replies, add CRM notesInternal writes, easy to fix
  • Pause an underperforming campaign variantStops harm; start it as Ask firstonce trust is earned

Ask me first

  • Send any email or replyReaches a real person, cannot be unsent
  • Launch a campaign or import contactsSends at scale
  • Change CRM deal stage or ownerAffects forecasts and other people
  • Spend money or change a subscriptionFinancial
  • Delete data or change permissionsIrreversible
Start everything in the right-hand lane as Ask first. Move one lane left only after a few weeks of being right.

Two cautions from the docs. An approval controls the proposed action; it does not undo work already done. And all Bots in an account share one cloud computer, so separate Bots are not a security boundary (docs).

Start everything in the bottom half as "Ask first." Move an action up one lane only after it has been right for a few weeks. Don't start any bot at full autonomy.

This has become an operating loop, not a conversation.

Zoom back out

Youdecisions only
First Mateone interface
GTM Second Mateowns the domain
Content Second Mate
Ops Second Mate
Consultantsjudgment
Specialists
Tools and other AIcapabilities
Routinesongoing responsibility
Always-on boxreach beyond Grok's computer
Work surfacesvisibility
Approval gates = controlled action
Every layer lit. Each one fixes the bottleneck the one above it exposed.

Go back to the map. First Mate: one interface. Second Mate: owns GTM. Consultants: judgment on demand. Tools and other AI: capabilities. Routines: responsibility. An always-on box: reach for tools outside Grok. Work surfaces: visibility. Approval gates: action you can trust.

None of these ideas is impossible in other tools. What Grok Bot changes is how easy each step is: a Bot is a name and a job, a group chat is a room, a routine is a skill plus a schedule, and approvals are built in. That ease is what lets you climb from one useful bot to an AI organization that finds and runs GTM work while you make the decisions.

Start today with one step: create the First Mate, paste the template, and hand it one messy request.

Sources

Bring the GTM job you need done. Build the run you can check.

Join the waitlist to hear when the next Legion cohort opens and what to bring to the first working session.