In Brazil, people are debating the 6x1 schedule: six days of work, one day off. I've been thinking about a different schedule. Here it's 24/7. And I'm not the one working 24 hours a day: it's the agentic software factory I built, as a hobby, in a little over a week.
This is the first chapter of a diary. I'll tell you what I'm building, how I got here, what went wrong along the way and what I learned from it. In the next chapters, each lesson gets its own piece, with the data and with the mistakes. For example, why the bottleneck isn't the code (chapter 4) and why git was made for humans (chapter 6).
How it started
On a Saturday night, I had a simple idea: what if I could open several AI sessions at once, each one taking care of part of the work? One writes the code, another writes the tests, another reviews. Maybe a few terminals and a script would do it.
It didn't stay in the terminals. That same night the idea became a desktop app with an engine that follows a workflow defined in a configuration file. Each step of the workflow is run by an AI agent (Claude, Codex or Grok), and the engine only decides which step comes next. It doesn't improvise.
That was the first important decision, and I only understood how big it was later. The part that decides "what comes next" is plain, predictable code that always does the same thing. AI only steps in where intelligence is needed: writing, testing, reviewing. That keeps the process auditable. When something goes wrong, I know exactly which step it was and why.
The workflow is the same one a development team would follow:
- Develop: an agent implements the task in an isolated copy of the project, without touching what already works.
- Test: another agent, one that didn't write the code, writes the tests.
- Review: a third one reviews, preferably with a different model from the one that wrote it. An AI reviewing another AI's work catches things the first one wouldn't see.
- Approve: the code only goes in if it passes the tests and the review. At the points where I want to decide, the workflow stops and waits for me. No automation goes over that approval.
That same night, the build steps, continuous integration and a test environment for every change went in too. I went to bed with a skeleton that already looked like a production line. It was the beginning of what is now called an agentic software factory: a software factory where every step is run by AI agents, and people decide what to build and approve what goes in.
Sunday: the first real run, and the first mistake
On Sunday I did what everyone does with a new toy: I ran it for real.
The first real run failed. A detail in how an agent's answer was validated brought the workflow down at the very end. The second one, with the full development workflow, worked, but it took almost 40 minutes, and the continuous integration step ran four times in a row before it passed. That was my first hint that the problem wouldn't be writing code. That hint became chapter 4, The bottleneck isn't the code, it's the tests.
Sunday was also when the supervisor was born: an AI that only steps in when a step fails. It investigates what happened and proposes how to get things moving again.
The first time it worked, it suggested fixing the code directly. It looked helpful, but I saw the risk: a supervisor that fixes code outside the workflow skips exactly the steps that make it safe, like testing, review and approval. I redefined its role on the spot. The supervisor takes care of the workflow, not the code. When the problem is in the code, it hands the task back to the workflow's agents, with the evidence and the hypotheses, and the code goes through every step again.
That rule became a principle of the project: automation never crosses a human approval, not even the supervisor.
Monday, the quiet day
Monday was the calmest day of the week. No runs, few requests. It was an architecture day: separating the engine from the interface, hardening the app's security and writing a context document that every agent reads before it starts. That document was born out of an emergency: the conversation with the AI got so long it was reaching its memory limit, so I asked for a summary of everything. The summary became the house manual.
Looking back, that quiet day is what made everything after it possible.
The day the tool started developing itself
On Tuesday, I opened several sessions in parallel, each one in an isolated copy of the project, and almost all of them started the same way: "run the app so I can see it".
In the middle of the afternoon, I opened a session just to jot down ideas in a text file. Fifteen minutes later, the ideas were already registered tasks, and the app got a backlog screen, where each task can be handed to a workflow. By midnight there were more than 50 registered tasks, almost all of them born from a conversation.
And then I started asking the tool itself to solve the tasks.
At around 4:40 pm, the first task was delivered by the tool itself. Almost in the same minute, I wrote something to Claude that now sounds like a prediction: that I wouldn't abandon it, I would keep using it, only, very soon, inside my own tool.
An hour later, at around 5 pm, the turn happened: the tool started generating more AI work than I did using AI directly. From then on, the curve never went back. In a few days, more than 90% of the AI work on the project was coming from the tool. My sessions changed their role: instead of writing code, I followed, diagnosed and decided. That turn has a chapter of its own: The day AI took the wheel from me (chapter 2).
Five agents, one file
On Tuesday night, excited, I put five tasks to run at the same time.
Each agent worked in its own copy of the project, without getting in the others' way. It worked very well until the end, when it was time to put everything together. Four of the five failed at the last step because of conflicts: two agents had changed the same part of a file. All the work had already been done and paid for, and it got stuck at the exit door.
That night taught me a lesson: parallelism is cheap to start and expensive to finish. I tell that story in chapter 6, Git was made for humans. The answers came in layers:
- Group related tasks into the same run, trading several conflicts for a single path.
- Merge the work locally, on my own computer, instead of depending on an external service for every change. That took the queues and the continuous integration minutes I didn't control out of the short cycle. The success rate of the runs jumped from about half to more than nine in ten.
- Plan ahead, instead of solving the conflict afterwards.
And it was that last layer that paved the way for the first night.
The first night
To leave the computer working on its own, I was missing someone to organize the work. So I created an AI Product Owner.
It reads the backlog and builds a roadmap in stages. For each group of tasks, it predicts which files will be touched, points out where conflicts may happen and puts the tasks that would fight over the same file in different stages. Conflict resolution moved from the moment the code is merged to the moment the plan is made. During the day, I write down the ideas. At night, the roadmap runs.
I remember the last message I sent that night: "cancel the monitor, I'm going to sleep and leave the PC running".
The roadmap ran in two parallel lanes. At 11:30 pm, the first two tasks. Close to 1 am, the next two. Then at 1:30, at 2:30 and at 3. The last run finished at four in the morning.
I woke up to 10 tasks delivered, tested and merged. None failed. Not a single message from me between midnight and seven in the morning. And among that night's tasks was one that put the supervisor itself on call. The tool was, literally, improving itself while I slept.
The machine only stopped at four because the roadmap ran out. There was no more planned work. That detail changed the way I think about the project: for 24/7, what matters most may be how many hours of well-defined work are ready when I go to bed. The behind-the-scenes of that night, hour by hour, is in chapter 3, I went to sleep and woke up to the work done.
The night everything almost went wrong
Not everything went smoothly. An hour before the first night began, I accidentally opened a second copy of the app, pointing to the same work folder. Both started writing to the same state at the same time. Runs that were alive were marked as interrupted, duplicate runs showed up and the roadmap stopped, pointing to cancelled tasks.
I asked the AI: "can't you fix it for me?".
And it could. Since all of the app's state lives in readable files, not in a closed database, the AI read the code that writes the roadmap, understood the format, made a backup, pointed each item to the right run and removed the pause. In two minutes the roadmap was running again, without losing any work. At that moment, I felt like I had superpowers.
The flaw and the superpower had the same origin: files have no locks. Two processes wrote to the same place. The definitive fix became a task that same day, and today the app guarantees a single instance per work folder. But the lesson stuck: a system the AI can read is a system the AI can fix. That idea, of keeping everything in files the AI can read, has its own chapter: A system AI can read is a system AI can fix (chapter 8).
The second night: non-stop
Before the second night, the tool spent the late afternoon reorganizing its own code so that several agents could work at the same time without bumping into each other. Screens split into their own files, lists with one entry per line, tables generated automatically. No change in behavior, only in format. It sounds small, but it's what made scaling possible.
And it scaled. On the second night the machine didn't stop. There were about 25 tasks, more than twice the first night, with up to 5 agents working at the same time. There were four conflicts, and all four were solved without me. At seven in the morning, a run was still going.
There was a single blocker, and I like it. In the middle of the night, a task reached a point where the workflow requires my approval. It sat there for more than two hours, waiting for me to wake up. That's exactly the right behavior: the machine works all night, but it doesn't make my decisions for me.
That same night, the credits of one of the AI providers ran out, and the reviewer switched to another model. Until then, I thought AI credits would never be the limit. With five agents in parallel, the consumption per hour more than doubles.
The moving bottleneck
The next day, I put four heavy runs in parallel, and my computer froze. I had to cancel everything.
That's when I noticed a pattern. The bottleneck never goes away, it moves. First it was the external continuous integration service. Then, the conflicts when merging the code. Then, the way the code was organized. Then, the planning. Now it was the machine itself: processor and memory. The answer was similar to a human team's: leave the heaviest tests for when a release is cut, not for every task.
The following night, the cost of the tests per run dropped to a fraction of what it was. And the bottleneck moved again: now it's the time the test suite takes to run. That story continues in chapter 4.
Two hours of mine, 22 hours of AI
When it finally sank in, it was shocking. I had no idea what it means to have agents working 24/7.
If I organize the workflows, the agents and the backlog well, I can interact with the system for two hours a day. In the first hour, I review what was done during the night and approve it or ask for changes. In the second, I prepare the work for the next 24 hours. The rest is AI working.
This project is what I do in my spare hours. And that's exactly what changed my mind: two hours a day of one person can turn into 24 hours of work, as long as there is well-defined work to do.
And there is never a shortage of work. Ask any company whether there's anything to do in tech, and the answer is always the same: there are a million things. What's missing is defining them, organizing them in a backlog and having a production line that delivers them with quality. The backlog is the agentic factory's fuel.
I haven't reached the two hours yet. On the second night, I still sent more than a hundred messages throughout the day, many of them to plan the night. But the path became clear: the limit is no longer my time.
What I've learned so far
The numbers catch the eye, but what was worth the most were the surprises. Almost all of them will become a chapter.
The bottleneck isn't writing code, it's testing (chapter 4). Adding up the time of each step, the tests took longer than the development. A good part of that time was an agent waiting for a command to finish, on conditions that were never met. Adjusting the tests cut the cost of that step to a fraction of what it was.
The most expensive token is the one of an AI waiting (chapter 5). The most expensive run of the project delivered nothing. An agent spent its time checking, again and again, whether a process had finished, blocked by a permission it didn't have. Waiting, checking status and comparing results is work for plain code. AI steps in when there's a decision to make.
Git was made for humans (chapter 6). People change a few lines at a time. Agents in parallel don't. Two agents that add a line at the end of the same list run into a conflict, even without any real contradiction between them. With several agents, even the format of the files becomes an architecture decision.
Refine first or react later? (chapter 7) The traditional way invests heavily at the start: refinement, prototypes, trying to predict every detail. When generating gets cheap, waking up to a finished proposal and saying "that's not it, it's like this" costs less than trying to describe everything in a vacuum. The working increment becomes the cheapest specification there is.
Whatever I ask for three times becomes automation. "Run the app so I can see it" became a step that produces evidence for every delivery, with screenshots and logs. "Register this as a task" became a workflow of its own. Even the voice dictation I used became a feature. My repeated requests were, almost always, the next piece of the product.
One week in numbers
- More than 200 commits since the first one, on a Saturday night.
- More than a hundred tasks registered, almost all of them born from a conversation.
- 10 tasks on the first night and about 25 on the second, with up to 5 agents at the same time.
- More than 90% of the AI work on the project coming from the tool itself.
- Zero human approvals crossed by automation.
Why I'm writing this
I'm not selling anything. It's a hobby project, and writing is the way I found to organize what I'm learning. I have the feeling that the way software is built is changing fast, and I want to record it from the inside, with real data, while it happens. Mistakes included.
In the next chapter, I'll tell you about the day AI took the wheel from me: how the turn happened, when the tool started doing more than I did, and what changes when you stop being the one who codes and become the one who decides.
And, in the last chapter, a conversation with the AI about the ship of Theseus that I never expected to have.