Principles for using AI effectively

Samuel Batista

Samuel Batista / July 18, 2026

11 min read • ––– views

I've been experimenting with running AI agents autonomously lately: kicking off loops, letting them work without me, and seeing how far they can get on their own. The experience has taught me a few principles that I think are useful to anyone working with AI. I'll keep the specifics of what I'm building to myself, but the principles stand on their own. Here they are.

1. Context is everything

The quality of what you get out of an AI is mostly determined by what you put in. Not just the prompt, but everything around it: the files it can see, the tools it can use, the instructions it carries, the feedback loops it has access to. Preparing that context is the real job of the AI engineer.

The formula I keep coming back to is simple:

Context + Model + Harness = Result

The harness is the environment the model runs in, whether that's a coding agent, a chat interface, or something you built yourself. A good harness feeds the model the right information at the right time so you don't have to do it by hand.

Taste matters here. Which model you pick, which harness you run it in, how you shape the context: each of these choices affects the quality of the output, and they compound. Most people focus all their attention on the model, but the model is only one term in the equation. In my experience, the context and the harness are where most results are won or lost, and they're the terms you have the most control over.

2. Build the machine that builds the machine

The mistake is treating the AI like a worker you hand tasks to. A task is usually a guess about how something should be built, and when it comes to the how, an AI with the right context often knows better than you do. Stop assuming you know how it should be built. Know what you want, and let the AIs figure out how to deliver it.

So instead of asking the AI to do the task, go one level up: ask it how you can build an AI that achieves the goal. Use the most powerful model you can get access to and have it help you design the system. The machine that builds the machine, as they say in manufacturing, where designing the factory is understood to be a harder and more valuable problem than designing the product.

I'll give away a small piece of how I do this. I've found it works well to have two kinds of agents: an architect and a worker. The human builds the architect, and the architect builds the worker. My job is to give the architect a clear goal and the right context. The architect's job is to figure out how the work should get done, and to produce the workers that do it.

If you don't know what you want yet, that's fine. Define a process that refines your ideas until you do. Vague goals produce vague results, but a clear goal handed to a well-built system can carry itself surprisingly far, even while you're doing something else.

3. Fork, rewind, and hand off your context

There's a catch to working at the goal level: goals outlive sessions. A task fits in one conversation. A goal takes many turns, with drafts, corrections, and dead ends along the way, and every one of them eats context.

Advanced use of AI usually involves forking and rewinding context. Forking means creating a copy of the agent at a point in time, so you can tinker with it safely and explore new directions without polluting the original. Rewinding means going back to an earlier point in the conversation and continuing from there. The value of rewinding isn't undoing your mistakes: it's cutting down the context the model is carrying without losing the work you've already done. The files you produced are still on disk; the tokens spent producing them don't need to stay in the window. Rewinding, forking, and compacting are your best friends when working with AI models on complex tasks.

While you're preparing the context for a task, you often discover it could be better, so you iterate on it. The catch is that all this iteration consumes the model's context window, and the context window is more fragile than advertised. Models may claim a one million token window, but in my experience they degrade noticeably beyond roughly 500,000 tokens. The degradation follows a pattern. The first thing you notice is that responses get slower. That part is just the mechanical cost of a long context, but it's your cue that you're deep in the window. The real degradation comes next: the model starts forgetting things mentioned earlier in the conversation, or gets lazy and does a worse job on its tasks. The last stage is outright hallucination, where it starts fabricating things.

So when you've spent a long session refining your context, don't keep grinding in a degraded session. Serialize what you've learned into a plan or handoff document, compact the session, and resume the work with a fresh context. Compacting means having your harness summarize the session and continue from that summary instead of the full history.

A pro tip on how I do this: near the end of a full session, I have the model write two things. First, a handoff document capturing the state of the work. Second, a handoff prompt that instructs the next agent to read that document, plus any extra context worth knowing from the session that didn't fit neatly into the document. Then I compact, run the handoff prompt, and the work resumes with a refreshed, more capable agent. The handoff document may sound redundant with compaction, but it isn't: the document captures what I've decided matters, instead of leaving that choice to an automatic summary.

4. Swap models to match the task

Not every task deserves the most powerful model. When the work is complicated or ambiguous, I reach for the strongest model available, either Claude Opus or Fable. These powerful models challenge my incorrect assumptions more effectively than weaker models do, and that pushback is worth a lot when I'm still figuring out what I want.

Once the hard part is done, like writing a detailed plan or implementing some tricky code, I swap to a weaker model like Claude Sonnet for the rest: commit messages, basic commands, and any other automation designed to save me time. A useful way to think about it: powerful models amplify your intelligence, weaker models save you time. Those are different jobs, and matching the model to the job keeps you fast and keeps your costs sane.

5. Let the AI triage your decisions

Here's a pattern I use constantly: when a step produces a lot of output, like a code review, a research report, or feedback from another agent, I don't read it all top to bottom. I have the AI read it for me and summarize the parts that actually need my attention. Then it walks me through the ambiguous ones one at a time, in an interactive question and answer session. For each decision I get its recommendation, the other options worth considering, and the reasoning behind them.

This isn't hypothetical. While editing this very post, an editorial review came back longer than I wanted to read. I had the AI triage it into what was worth acting on and what wasn't, and then ran this prompt to make the final calls:

Run through each of these feedback points one by one and help me make a decision, include your recommendation and other options for me to consider

This keeps you in the loop where you matter, on the judgment calls, and out of it where you don't. The AI does the reading; you do the deciding. It's principle 2 applied to information instead of construction: know what you want, and let the AI filter everything else down to the decisions only you can make.

The obvious objection is: how do you know the output is any good? That's the art of the craft. Judging the value of what an AI produces is still the job of the engineer. You can ask an AI to evaluate its own work, but models tend to be blind to their own biases and shortcomings, just like humans. The state of the art right now is to have different frontier models check each other: have one lab's best model build the feature, and another lab's best model review it. I won't be more prescriptive than that, because this is a nascent technology and the techniques are still developing. I'll leave it as an exercise for the reader to come up with novel ways of judging the value of their AI's outputs.

6. You can't be taught this, you have to tinker

The last principle is the uncomfortable one: reading this post won't make you good at AI. AI is not easy to teach. I can share these principles, but the only way to build real mastery is to tinker. Practice makes perfect, and there's no shortcut around the practice.

In fact, the thing that changed how I work the most wasn't any single technique. It was going all in: using AI for everything, even trivial things, purely to build the muscle memory of working with these tools. The most visible change is that a lot of my programming now happens by speaking to my computer, using Handy. I still type. I still review and edit code in VS Code. But talking to my computer? That's new.

Part of that practice should be using a variety of models. The more models you use, the better your intuition gets for what each one is good at, where it falls short, and what the right size of task is to hand it.

One thing that surprised me: a model's strengths and weaknesses seem to follow the company that makes it, not the individual release. A new version from the same lab tends to feel like a sharper version of the same model, with similar instincts and similar blind spots. Once you've internalized those tendencies, you can pick the right model for the job almost by feel. But you only get there by using them, a lot, on real work.

So take these principles for what they are: a head start, not a substitute for practice.

Addendum: run your AI during off-peak hours

One more observation worth sharing. I've been using Anthropic's models for a while now and have mostly specialized in them. I find them to be the most powerful, and frankly, the most usable as an engineering partner. But I've noticed a pattern: during weekdays, at peak hours, I run into more errors and slower responses. On weekends and late at night, the same models are snappy and reliable.

The strangest thing I've seen: I once caught Claude outputting words in Japanese mid-response and correcting itself afterward. I honestly don't know what caused it. It could have been an ordinary glitch, or an infrastructure bug: Anthropic has documented one with nearly identical symptoms. They state that they never reduce model quality due to demand, and I take them at their word. I'm sharing it because it happened, not because I can explain it.

What I can say with confidence is that the slowdowns and errors at peak hours are real, and they add up when you're running agents on long tasks. Off-peak, in my experience, you get better response times, fewer failed runs, and agents that can chew on your problems for longer without interruption. If you're running agents autonomously anyway, this is easy to take advantage of: let them work while everyone else is asleep.


Written with Claude (Fable 5) from my own stream-of-consciousness notes. Practicing what I preach: I provided the context, picked the model, and let it do the writing.