We built our project manager by talking to it
The technical project manager of our product team is an AIsuru agent: it writes the Linear tasks, watches cycle capacity, looks for the origin of bugs in the code and sends the reports. We built it by talking to it, and in the end it wrote its own prompt.
For a few weeks now the technical project manager of our product team has been an AIsuru agent. It turns a sentence into a well-written Linear task, watches the load of each cycle and speaks up when one is about to overflow, digs into triage tickets looking for the cause in the code, and sends the team the weekly and end-of-cycle reports. It proposes; it does not decide.
The interesting part is not so much what it does, but how it was built: we taught it by talking to it, and in the end it wrote its own prompt.
The problem, before the agent
A team of four does not have a full-time project manager, and does not want one. It has the same problems as a large team, though: tasks written in a hurry that nobody can test, a cycle that fills up without anyone noticing until it is too late, support tickets sitting still because assessing each one costs half an hour, and the fact that people who do not write code have no idea what is going on until someone tells them out loud.
These are all activities that take context and consistency rather than extraordinary intelligence. That is exactly the kind of work an agent can take on, provided it is taught properly and given access to the same tools we use.
What it does, and what it is not allowed to do
The agent lives on a single Linear team and a single project, and any request pointing elsewhere is declared out of scope. Inside that scope it does four things.
It writes the tasks. A developer says a sentence — "I'm fixing a bug in the login" — and the agent first checks whether something similar already exists, then writes title, description, priority, assignee and cycle following a template that always includes the closing conditions and the test cases for whoever does QA. The person who is working does not have to stop and write: the agent writes.
It watches cycle capacity. We do not use estimates, so the agent builds its own per-task weight from the priority and compares it with the median of what the team actually completed over the last three closed cycles. When one more task pushes a cycle over the line, it says so immediately and proposes what could come out, showing the arithmetic and declaring every time that this is its own estimate, not a figure from Linear.
It assesses triage tickets. Three times over — daily, weekly, monthly — it looks at the tickets coming in from support, classifies them by urgency and, before classifying them, searches the code for the likely origin of the problem: repository, file, function, and when the fix is genuinely visible even a snippet, in a comment on the ticket. If it finds nothing plausible it says so in one line instead of inventing a lead, because a wrong lead costs a developer more than no lead at all.
It sends the reports. Weekly on Monday, end-of-cycle on the first day of cooldown, plus alerts whenever triage turns up something urgent. They are written on two registers: a summary anyone can read at the top, the technical detail underneath.
The constraint that governs everything else is that the agent proposes and people decide. There is an explicit table of what it may do without asking, split between conversation and automated run, and the automated runs never move someone else's work, never lower a priority and never close anything. There is also a list of things it must never do: delete, write to addresses outside the list, carry a customer's data into a report, and — the one that matters most of all — report an action as done without having re-read it to confirm.
The tools
The agent talks to four systems, all through MCP, the protocol an AIsuru agent uses to reach external tools.
Linear is the system of record: it reads and writes issues, cycles, comments, labels and states. Outlook sends the reports. The AIsuru Scheduler wakes it up at the right times. And then there is the piece that, in our view, makes the difference between this and ordinary project-management automation: an MCP over our own codebase, read-only, exposing search, reading of line ranges and git history across the product repositories. No shell, no writes, no pull. That is what lets the agent answer "this bug probably starts here" instead of merely sorting tickets.
How we built it
The analysis came first, written by a human. As with any project, you start from the requirements. They were written by the CTO, who owns the development process: every responsibility described one by one, the workflows, the states, what happens when a developer is in a hurry, what each report has to contain. A discursive document, not a schema.
The device that proved more effective than anything else in that document was the scenarios: short four- or five-line dialogues between a developer and the agent, written the way they would actually have gone.
— I'm fixing a bug in the login page.
— Noted. I'll add it to the current cycle and move it to In Progress. Any details for me?
— No, I'm in a hurry, I'll add them later.
— Created, assigned to you, Urgent priority. I'm sending the CTO a proposal of what to move to the next cycle, because this one now overflows.
Four lines like these pin down the behaviour better than a hundred words of description: they carry the tone, they say how much the agent may infer on its own, they say where it stops and who it warns. And above all they resist ambiguity — if a scenario is wrong you notice it as you read it, whereas a paragraph of requirements can look right and not be. Wherever we wrote scenarios, the final prompt came out correct on the first attempt; where we only had prose, we had to clarify afterwards.
Then the translation into a prompt, done by an LLM. The document was handed to Claude Opus 5 with the job of turning it into operating instructions. LLMs are very good at writing prompts for other LLMs: given a good description to start from, what came out was a document in two parts — one for whoever configures the agent, one meant to become its runtime instructions — with the definitions made measurable, the autonomy rules put into a table and the templates ready to use.
With one deliberate limitation: that model did not know AIsuru. It knew how to write a prompt, not which tools actually exist on the platform or how they are configured. The result was excellent as structure and inaccurate as context.
Then the agent, which finished writing itself. We created the agent on AIsuru and enabled the MCPs it needed. Then, for the construction phase, we also gave it the Vibe Coder.
The Vibe Coder makes a great deal possible, but the part we needed here is twofold: on one side it puts the AIsuru documentation at the agent's disposal, along with the knowledge of how MCPs are used and how agents are configured on the platform; on the other it grants it builder permissions, that is, the ability to rewrite its own instructions and connect itself to its own tools. Put together, they mean the agent is not something you configure from the outside: it is a counterpart that can work on itself while you explain what it has to do.
From there, building it became a conversation. We told it we were going to hand over two analysis documents so that it could update its prompt and enable the tools, and it answered that it was ready. We gave it the analysis document as the primary source, with the instruction to read it and say whether it was clear, and it came back with the right questions — the missing data, the ambiguities, the things the document took for granted. We cleared them up there, in the chat.
Then we gave it the second document, the one generated by the LLM, explaining what it was: a proposal written by someone who did not know the platform, to be taken for what it was. It read it, checked a couple of things directly on Linear rather than answering blind, and told us which parts were worth keeping as they were — the capacity weights, idempotency through labels, the report templates — and which had to be rewritten because they did not match how the tools on AIsuru actually work. Then, with builder permissions, it wrote its own prompt and configured the six scheduler jobs by itself.
That is the step that struck us most: we did not transfer the knowledge of the platform to it. It already had it, and it was that knowledge that corrected the work of the larger model.
What came out of it
Today the agent runs on a prompt of roughly thirty-two thousand characters, sixteen configured functions — seven MCPs, each with its own pair of tools to list and to execute — and six scheduled jobs: daily, weekly and monthly triage, the weekly report, the end-of-cycle report and the planning reminder.
One detail says a lot about how it is put together: the jobs contain no logic. The payload the scheduler wakes it with is a block of context identical for all of them — you are running a task on your own, nobody is in the chat, do not ask for confirmations — followed by a single line, for instance run: end-of-cycle-report. All of the behaviour lives in the prompt. When we had to change the recipients and the language of the reports, we did not touch a single job.
One real run, the end-of-cycle one, to give the measure: seven minutes from invocation to report sent, in which it classified the triage tickets still open, drafted the release notes for a task that had just moved into communication, reconstructed the state of the closed cycle and sent the email to the team, with the list of what had not been completed and a proposal for each item. No task moved: proposals only, as designed.

Which is why idempotency is not a detail: the agent has no memory between one run and the next, so it does not remember what it has already done. The state lives on Linear, in the form of labels — an assessed ticket carries one, a ticket waiting for an answer from support carries another — and every run queries the labels, never its own recollection.
Three things we take away
The analysis document matters more than the prompt. The prompt is a translation: if the analysis is precise, a model does the translation in minutes. The time spent writing good requirements is the only time that cannot be compressed.
A general-purpose model writes excellent instructions and poor context. Knowing how to write an effective prompt and knowing how your platform works are two different competences. Keeping them apart — the model for the structure, the agent for the context — worked better than asking one of them for everything.
The right limit is not technical, it is about autonomy. The hard question was never what the agent is capable of doing, but what it is allowed to do without asking. Having answered that in a table, rather than case by case, is what lets us sleep while it runs on its own at half past seven in the morning.