On this page
The fix wasn't a better bot. It was a management layer. Grok Bot makes that layer easy to build.
By the end you'll have a structure you can copy: one bot you talk to, one bot that owns GTM, a bench of experts, recurring jobs, and clear rules for what the system may do without asking.
The map
Here is the whole model before any detail:
- You talk to one bot: the First Mate.
- The First Mate hands work to Second Mates. Each one owns a domain, like GTM.
- Second Mates pull in consultants for judgment and tools for execution.
- Routines give each owner jobs that run without a prompt.
- Work surfaces, like your CRM and campaign platform, let the system see real state.
We start with the part you can set up this afternoon and add one layer per section. Each layer fixes the bottleneck the previous one exposes.
Layer 1: The First Mate is your only front door
The simplest change has the biggest payoff. Pick one bot and make it the only one you talk to by default. Its job is not to know everything. Its job is to know who does what and send work there.
Without it, you are the router. You remember which bot handles research, which one writes, and which one has the campaign context. That routing is the context switching you were trying to hand off.
Setup: create the First Mate
In Grok Bot, you create a Bot by giving it a short name, one primary job, and a description of how it works (docs). For the First Mate:
- Name: First Mate.
- Primary job: take every request from me, decide who owns it, delegate, and report back.
- Paste in the operating template below and edit it to your business.
- Test it with one messy request (example further down) before you add any other bots.
ROLE and ABOUT ME: the job, the business, the quarter's goals, how you like updates.
One line per Second Mate. If nobody owns it, say so and propose an owner.
One domain, one owner. Split mixed requests. Anything that sends, spends, publishes, or deletes comes to you.
What you asked for, who took each part, what is done, waiting, or needs a decision.
Example First Mate operating template
ROLE
You are my First Mate. I talk to you, not to the other bots.
Your job is coordination: route work to the right owner, track it,
and report back. You do not do domain work yourself if an owner exists.
ABOUT ME
- Business: [one line on what we sell and to whom]
- This quarter's goals: [2-3 goals with numbers]
- How I like updates: short, decision first, evidence after.
WHO DOES WHAT
- GTM Second Mate: campaigns, lists, messaging, replies, pipeline.
- Content Second Mate: newsletter, social, site copy.
- Ops Second Mate: billing, tooling, admin.
If nobody owns a request, tell me and propose an owner. Do not guess.
DELEGATION RULES
1. One domain, one owner. Never give the same job to two bots.
2. Split a mixed request into parts and send each part to its owner.
3. Pass the owner the goal, the deadline, and what done looks like.
4. Anything that sends, spends, publishes, or deletes comes to me first.
REPORT FORMAT
- What I asked for
- Who took each part
- What is done, what is waiting, what needs my decisionKun Chen's open-source firstmate is a useful companion here. It is an "agent distro" built on the same idea for coding agents: you talk to one agent and it runs the crew, each worker in its own copy of the code. Grok is listed among its supported harnesses. Read its AGENTS.md for templates you can adapt to your First Mate.
The messy request test
Give your First Mate something that crosses three owners:
"The fintech campaign replies dropped this week, the newsletter is due Thursday, and I think we're double paying for an enrichment tool. Sort it out."
A working First Mate does not try to answer all of it. It sends the reply drop to GTM, the newsletter to Content, and the billing question to Ops, then comes back with one status list.
- “replies dropped”GTM Second Mate
- “newsletter Thursday”Content Second Mate
- “double paying”Ops Second Mate
- GTM: investigating reply drop
- Content: draft due Thursday
- Ops: needs your decision
You do not need to remember which bot handles which task. You talk to one management agent.
Layer 2: Second Mates own domains, not questions
Once the First Mate routes work, the next question is where it lands. A Second Mate is a bot that owns a domain: its projects, context, tools, and specialists. GTM is the best one to build first, because GTM work is a chain of steps that all need the same context.
The rule is one domain, one owner. If three bots all do "research and strategy," they overlap, contradict each other, and you end up judging between them. Give each domain one owner and the overlap disappears.
Example GTM Second Mate operating template
ROLE
You are the GTM Second Mate. You own outbound and pipeline for [company].
You take work from the First Mate and report back to it, not to me,
unless a decision needs me.
WHAT YOU OWN
- Who we sell to: [segments, titles, company size, disqualifiers]
- Offer and messaging: [current offer, proof points, banned claims]
- Active campaigns: [name, segment, status, goal]
- Experiments: [what we're testing, success threshold, end date]
- Pipeline: [where deals live, what "stalled" means, e.g. no touch in 10 days]
YOUR TOOLS AND SPECIALISTS
- List building and enrichment: [tool / connector]
- Campaign platform: [tool / connector]
- CRM: [tool / connector]
- Consultants you may call: [see bench below]
OPERATING RULES
1. Start every project from a written brief: goal, segment, volume, deadline.
2. Show your work: counts at each step, and what you dropped and why.
3. Draft anything outbound. Never send, launch, or import without my approval.
4. When a metric moves more than [X]% week over week, investigate before reporting.
DAILY ANSWER TO "WHAT SHOULD WE FOCUS ON TODAY?"
Top 3 items, each with: the evidence, the owner, the next step.Ask it "what should we focus on today?" If the answer draws on your campaigns, experiments, and pipeline rather than generic advice, the Second Mate is working.
A bot can own a domain, not just answer questions about it.
Layer 3: Run one real GTM project end to end
This is where the hierarchy proves itself. Give the GTM Second Mate one concrete outcome and let it coordinate every step. The leverage is not writing emails faster. It is removing the human handoffs between the steps.
An example brief: "Build a campaign for Series A fintech companies with 20 to 200 employees whose head of sales was hired in the last six months. Target 500 verified contacts. Draft three messaging angles. Have it ready for my review Friday."
- Brief1 goal
- Researchsegment notes
- Accountsexample: 1,200
- Contactsexample: 2,100
- Qualificationexample: 640 qualified
- Messaging3 angles
- Campaign draftexample: 500 verified
- Your approvalbefore launch
- Launchafter yes
- Monitoringroutines
The Second Mate runs the chain: research on the segment, account sourcing, contact finding, email validation, qualification against your rules, messaging, and a campaign draft. At each step it reports counts and what it dropped. You see one deliverable, not seven tool sessions.
For the list-building steps, the Second Mate can hand work to a specialist bot. Clay Desk by Mitchell is one I built for this, and you can use it. Its page says it runs Clay work end to end: it builds tables, runs enrichments and functions, and uses a real browser when a step only exists in the Clay app. It shows a cost estimate and asks before spending credits, which is the approval boundary this whole article builds toward.
It does not need to be the best tool at every step. Some steps run through connectors, some through other tools. Its value is keeping one goal and one context moving through all of them.
The leverage is coordinating the workflow, not doing any one task slightly faster.
Layer 4: Consultants for judgment calls
Execution problems have answers. Judgment problems have tradeoffs. When the campaign above gets a 1% reply rate, "run the next step" is not the fix. You need people who see the problem differently.
Consultants are bots or saved contexts that represent a framework, a discipline, or a school of thought. Grok Bot lets Bots share context in group chats and pass work to each other (docs), so the Second Mate can open a temporary room with three or four of them.
Example consultant bench for GTM
| Consultant | Lens | The question it always asks |
|---|---|---|
| Offer skeptic | Value of the offer to a cold reader | Why would a stranger say yes to this in one email? |
| List purist | Targeting and data quality | Are these the people who can buy, or people who can only recommend? |
| Deliverability hawk | Infrastructure and inbox placement | Are these emails reaching the inbox at all? |
| Buyer proxy | The prospect's point of view | What would make me delete this in two seconds? |
You can also distill a real expert into a consultant. Distill Anyone is a Grok Bot built for this. Its page says it builds a talkable companion from someone's public YouTube transcripts, adds topic depth over time, and refreshes daily.
The useful part is disagreement. Tell each consultant to argue from its lens and attack the others' reasoning. A room where everyone agrees is one model talking to itself four times.
- Offer skeptic
A stranger has no reason to say yes to this offer in one email.
- List purist
Half this list can only recommend, not buy. Fix targeting first.
- Deliverability hawk
Check inbox placement before you touch the copy.
- Buyer proxy
I'd delete this in two seconds. The first line is about you, not me.
- Likely cause
- List: too many recommenders, not buyers
- Next test
- Re-cut to budget owners, same copy
- Owner
- GTM Second Mate, report Friday
Then the Second Mate synthesizes. It does not hand you four opinions. It hands you a decision: the likely cause, the next test, and who runs it.
You can assemble expertise around a problem on demand instead of asking one model to role-play everything.
Layer 5: Grok as the front door to your other AI
You probably already pay for other AI: a coding agent, a research tool, another model you like for writing. The usual pattern is switching between chat windows. The better pattern is treating each one as a capability the hierarchy can call.
Each Grok Bot runs on a cloud computer with a browser, filesystem, and terminal, and uses connectors where available and computer use for everything else (docs). That is what makes this possible: if a tool runs in a browser or a terminal, a Bot can be taught to use it.
In the fintech project, say you want a deeper market scan than a quick search. The Second Mate hands that to a research tool, waits for the result, and folds it into the brief. You never open the other app.
The manager doesn't need to out-research or out-write the specialist tools. It needs to know which one to hand the task to. This is a workflow pattern, not a named Grok Bot feature. How it works depends on each tool's login and terms.
Grok becomes the interface to your other intelligence instead of competing with it.
Layer 5.5: Keep context clean with one Code Mode MCP instead of fifty plugins
Once Grok is the front door to your other tools, the obvious move is to plug everything in. CRM, campaign platform, enrichment, inbox, docs. Each one ships as a plugin or MCP server, and each one adds its tool definitions to the model's context before you type a word.
That is where it breaks. Anthropic's engineering team describes agents connected to hundreds or thousands of tools having to process hundreds of thousands of tokens of tool definitions before reading the request (Anthropic). Every tool result also flows back through the context, even when the agent only needs one field from it. The window fills with schemas and raw data, and the answers get worse. That is context rot.
The fix is called Code Mode. Instead of loading every tool, you give the agent one tool that runs code. The agent looks up the method it needs, writes a few lines that call it, and only the filtered result comes back into context.
Cloudflare introduced it: their Code Mode turns MCP tools into a TypeScript API and has the agent write code against it, because models have seen far more real code than tool-call syntax (Cloudflare). Anthropic's version of the same idea cut one workflow from 150,000 tokens to 2,000, a 98.7% reduction (Anthropic).
- crm.search
- crm.update
- crm.notes
- campaigns.list
- campaigns.stats
- campaigns.pause
- leads.enrich
- leads.verify
- inbox.search
- inbox.reply
- docs.read
- docs.write
- calendar.find
- billing.invoices
- sheets.append
- slack.post
- +34 more
50 tool definitions in context before the first word of work
- codemode
- describe("crm.search")
- result: 3 fields x 12 rows
Methods looked up on demand, results filtered in code
How we run it
Our integrations sit behind one gateway that exposes one codemode tool. The agent never sees fifty schemas. When it needs a method, it calls codemode.describe("<connector>.<method>") to read that one signature, then writes a short script that calls it. Anything that writes to the outside world pauses until a human approves it.
Here is what the agent's code looks like. This is illustrative, not a real API:
// Illustrative only. The agent looks up one method, then calls it.
const spec = await codemode.describe("crm.searchDeals");
const deals = await crm.searchDeals({ stage: "proposal", updatedWithinDays: 14 });
// Filter before anything returns to the model's context.
return deals
.filter(d => d.amount > 20000)
.map(d => ({ name: d.name, owner: d.owner, lastTouch: d.lastActivity }));The model reads three fields per deal instead of the full CRM payload. The raw data never touches the context.
Set it up
- Pick a gateway approach - Use Cloudflare's Code Mode, or an MCP gateway that puts many connectors behind one endpoint. Either works if the agent writes code instead of calling fifty tools.
- Register your connectors behind it - CRM, campaign platform, enrichment, inbox. Add them to the gateway, not to the agent.
- Expose only two or three tools - A search or describe tool for finding methods on demand, and an execute tool that runs the code. Nothing else goes into the agent's context.
- Keep the keys in the vault - The gateway pulls API keys from your secrets manager at run time, the same Infisical setup from Layer 7. No key sits in a prompt or in the agent's code.
- Put an approval gate on writes - Reads run directly. Anything that sends, spends, updates or deletes pauses until you approve it. This is the same boundary as Layer 10.
- Connect Grok to it - The xAI API supports remote MCP servers by URL (docs). Grok Bot's docs only describe connectors installed from its Marketplace (docs), and do not document adding your own MCP server. So treat this step as workflow-specific: reach the gateway through the always-on box in Layer 7, or through the API.
One tool in context instead of fifty. The Bot spends its window on your work, not on reading tool manuals.
Layer 6: Routines turn prompts into responsibilities
Everything so far still waits for you to ask. Routines change that. A routine tells one Bot when to run a workflow, on a schedule or, where supported, after an event, and it usually runs a saved skill (docs).
Think "this is your job every morning," not "run this prompt every morning." The difference is ownership: the Bot owns the outcome and decides what is worth reporting.
The routine I'd set up first isn't a GTM one. It's a bookmark reviewer. It browses alongside you, learns what you engage with, and turns it into drafts for your next posts, then asks what you think so every piece carries your opinion. It is the fastest way to feel a Bot working like you.
Example routine definitions
Bookmark reviewer and builder
- Owner
- Content Second Mate
- Trigger
- Daily, plus whenever I bookmark or save a post
- Job
- Browse what I bookmarked, liked, and engaged with. Learn what I react to and why. Turn the strongest items into post ideas and drafts that use them as inspiration, and ask for my opinion on each so the next post carries my take, not a summary.
- Report
- Top 3 ideas with the source post, the angle, and one question for me.
- Approval
- Publishing anything.
Morning campaign review
- Owner
- GTM Second Mate
- Trigger
- Weekdays 7:30 AM
- Job
- Check every active campaign against last week. Flag any reply rate or bounce rate that moved more than 25%. Investigate before flagging.
- Report
- Only if something moved. Otherwise one line: "All campaigns normal."
- Approval
- Pausing or editing a campaign.
Reply triage
- Owner
- GTM Second Mate
- Trigger
- Every 2 hours during business days
- Job
- Sort new replies into interested, not now, wrong person, unsubscribe. Draft responses for interested and wrong-person replies.
- Report
- Count by category, plus drafts waiting for review.
- Approval
- Sending any reply.
Stalled pipeline check
- Owner
- GTM Second Mate
- Trigger
- Mondays 8:00 AM
- Job
- Find deals with no activity in 10+ days. Suggest a next step for each.
- Report
- Top 5 stalled deals by value.
- Approval
- Any CRM stage change.
Weekly experiment review
- Owner
- GTM Second Mate, reporting to First Mate
- Trigger
- Fridays 3:00 PM
- Job
- For each running experiment, compare to its success threshold. Recommend keep, kill, or extend.
- Report
- Always.
- Approval
- Ending an experiment or starting a new one.
ROUTINE: Bookmark reviewer and builder
Owner: Content Second Mate
Trigger: Daily, plus whenever I bookmark or save a post
Job: Browse what I bookmarked, liked, and engaged with. Learn what I react to
and why. Turn the strongest items into post ideas and drafts that use
them as inspiration, and ask for my opinion on each so the next post
carries my take, not a summary.
Report: Top 3 ideas with the source post, the angle, and one question for me.
Approval: Publishing anything.
ROUTINE: Morning campaign review
Owner: GTM Second Mate
Trigger: Weekdays 7:30 AM
Job: Check every active campaign against last week. Flag any reply rate
or bounce rate that moved more than 25%. Investigate before flagging.
Report: Only if something moved. Otherwise one line: "All campaigns normal."
Approval: Pausing or editing a campaign.
ROUTINE: Reply triage
Owner: GTM Second Mate
Trigger: Every 2 hours during business days
Job: Sort new replies into interested, not now, wrong person, unsubscribe.
Draft responses for interested and wrong-person replies.
Report: Count by category, plus drafts waiting for review.
Approval: Sending any reply.
ROUTINE: Stalled pipeline check
Owner: GTM Second Mate
Trigger: Mondays 8:00 AM
Job: Find deals with no activity in 10+ days. Suggest a next step for each.
Report: Top 5 stalled deals by value.
Approval: Any CRM stage change.
ROUTINE: Weekly experiment review
Owner: GTM Second Mate, reporting to First Mate
Trigger: Fridays 3:00 PM
Job: For each running experiment, compare to its success threshold.
Recommend keep, kill, or extend.
Report: Always.
Approval: Ending an experiment or starting a new one.Per current docs, a Bot can own up to 50 routines, and you should test a routine before enabling it, because "a test run performs real work" (docs).
A recurring responsibility is different from a prompt you manually repeat.
Compute, not chat messages
Once routines run, think about usage differently. The one-line "all campaigns normal" message might sit on top of dozens of steps: several Bots, connector calls, browser sessions, another model, and retries.
Fintech campaign reply rate down 40%. Likely cause: new segment. Recommend pausing variant B.
- 3 Bots involved
- Connector calls: campaign platform and CRM
- Browser steps
- 1 external research call
- 2 consultant turns
- 1 retry
So budget for work performed, not messages. Grok Bot usage resets weekly (docs), which makes a weekly budget the natural unit. Start with fewer routines, watch what they consume, then add more. Plan tiers and limits change, so check Plans and billing before relying on any number.
One small output can represent a lot of work, and that is what you are paying for.
Layer 7: An always-on environment for tools outside Grok
First, what you do not need. Grok Bot work runs on its cloud computer, and the docs say closing the app, your laptop, or your phone does not stop a background turn or a routine (docs). You do not need a server to keep native routines alive.
A VPS (a rented computer that stays on in a data center) matters for a different case: tools that live on your own machine. A command-line tool you configured locally, a script, a browser session on your laptop, or software you want the system to reach at 3 AM. Move those to an always-on computer, and they stop depending on whether your laptop is open.
Keep it simple: a VPS is an always-on computer for capabilities that would otherwise live on your laptop. This is a workflow choice. How a Bot reaches that computer depends on your setup, for example a terminal login or a custom MCP server (a small service that exposes a tool to the Bot). It is not a documented default.
Once more than one agent needs the same API keys, keep them in one secrets manager instead of pasting them into each bot's instructions or leaving them in files on the server. We use Infisical: every agent, on the VPS or anywhere else, pulls the keys it needs at run time from the same place. You rotate a key once and every agent picks it up, and no key ever sits in a prompt or a chat log.
External tools do not have to depend on the laptop you happen to be using.
Layer 8: Let Grok see where the work lives
Until now, the system mostly knows what you tell it. That caps its usefulness, because you become the person copying the state of the business into chat.
Connect the places where GTM state actually lives: CRM, inbox, campaign platform, analytics, a database like Supabase, or your project tracker. Grok Bot ships connectors for tools like Gmail, Google Calendar, Google Drive, Outlook, Teams, and Salesforce, and you can add more through the Marketplace. That list comes from third-party guides, so confirm the current one in the app.
Connect only what a workflow needs. The docs say it directly: "Connect only the tools a workflow needs" (docs).
Now the GTM Second Mate can see what you see: unhandled replies, a campaign whose reply rate fell, contacts missing enrichment, deals aging in one stage.
The AI no longer needs you to copy the state of the business into chat.
Layer 9: Let Grok find the work
This is where the direction of work reverses. Once the system can see real state, and routines check it, it can notice problems before you do.
- 1
You notice, you ask AI.
- 2
AI inspects when asked.
- 3
AI inspects on a routine.
- 4
AI notices, investigates, brings it to you.
- 5
AI notices, delegates, prepares the action.
- 6
AI executes permitted actions, updates the system, reports back.
Asking a bot to check something is step 2. Routines get you to step 3. Step 4 is the jump: the Second Mate spots that the fintech campaign's reply rate halved, checks whether bounces rose, sees that one new segment accounts for the drop, and arrives with the cause and a recommendation.
- You
- notice problem
- open bot
- paste context
- ask
- read answer
- decide
- Signal in campaign platform
- GTM Second Mate notices
- investigates
- First Mate
- You receive one decision card
The AI can notice the problem and come to you.
Layer 10: Controlled execution closes the loop
The last layer lets the hierarchy act, inside limits you set. The full chain looks like this: a signal appears, the Second Mate notices, investigates, pulls in a consultant if the call needs judgment, prepares the next step, takes the actions it is allowed to take, updates the shared surface, and reports up.
Grok Bot has the controls for this. You can set limits in the request itself, and admins can set auto-review rules of "Ask first" or "Allow automatically," where "Ask first" wins when both match (docs). The docs recommend approval for sending, purchasing, deleting, publishing, or changing production systems.
Approval boundaries for a GTM Second Mate
Do it
- Read CRM, inbox, campaign statsRead-only, no risk
- Research accounts, find and validate contactsReversible, internal
- Draft emails, replies, and campaign copyNothing leaves the building
Do it, tell me after
- Tag replies, add CRM notesInternal writes, easy to fix
- Pause an underperforming campaign variantStops harm; start it as Ask firstonce trust is earned
Ask me first
- Send any email or replyReaches a real person, cannot be unsent
- Launch a campaign or import contactsSends at scale
- Change CRM deal stage or ownerAffects forecasts and other people
- Spend money or change a subscriptionFinancial
- Delete data or change permissionsIrreversible
Two cautions from the docs. An approval controls the proposed action; it does not undo work already done. And all Bots in an account share one cloud computer, so separate Bots are not a security boundary (docs).
Start everything in the bottom half as "Ask first." Move an action up one lane only after it has been right for a few weeks. Don't start any bot at full autonomy.
This has become an operating loop, not a conversation.
Zoom back out
Go back to the map. First Mate: one interface. Second Mate: owns GTM. Consultants: judgment on demand. Tools and other AI: capabilities. Routines: responsibility. An always-on box: reach for tools outside Grok. Work surfaces: visibility. Approval gates: action you can trust.
None of these ideas is impossible in other tools. What Grok Bot changes is how easy each step is: a Bot is a name and a job, a group chat is a room, a routine is a skill plus a schedule, and approvals are built in. That ease is what lets you climb from one useful bot to an AI organization that finds and runs GTM work while you make the decisions.
Start today with one step: create the First Mate, paste the template, and hand it one messy request.
Sources
- Grok Bot overview: docs.x.ai/grok-bot/overview
- Get started: docs.x.ai/grok-bot/get-started
- Skills and routines: docs.x.ai/grok-bot/skills-routines-and-automations
- Approvals, security, and privacy: docs.x.ai/grok-bot/approvals-security-and-privacy
- Cloudflare, Code Mode: blog.cloudflare.com/code-mode
- Anthropic, Code execution with MCP: www.anthropic.com/engineering/code-execution-with-mcp
- xAI, Remote MCP tools: docs.x.ai/developers/tools/remote-mcp
- Grok Bot, computer and apps: docs.x.ai/grok-bot/computer-and-apps
- Plans and billing: cursor.com/help/grok-bot/plans
- Clay Desk by Mitchell (Grok Bot): x.ai/bot/-TgkSGmlHRwY1GCQMzLpD
- Kun Chen, firstmate: github.com/kunchenguid/firstmate
- Secondary (connector list, SuperGrok tier): www.wrightmode.com/blog/grok-bot-setup-guide, mem0.ai/blog/grok-bot-guide

