2026-08-14
My coding setup as of August 2026
My AI agent workflow as of August 2026: Hermes as orchestrator, OMP for implementation, and open-source models at a fraction of the cost.
The intro
This post was originally going to be called “My coding setup”, but after thinking about it, I decided the title needs a timestamp. My workflow has nothing to do with what it looked like in June 2026, April 2026, or December 2025.
Since the start of the year, my main harness for everything has been OmO, running on OpenCode.
At the end of May, something changed.
I had heard and read about Hermes Agent, but I had never really tried it properly. Finally, I gave it a shot and it turned out to be an incredible discovery.
Before getting into the details, here’s an overview of my stack right now. My knowledge base lives in: Obsidian, Gmail, WhatsApp and GitHub, with the occasional extra source, but those are the main ones. Obsidian holds all my markdown files: SOPs (Standard Operating Procedures), meeting notes, research, X posts, tools, articles and a long etcetera, which deserves a post of its own. Gmail and WhatsApp handle client communication, and GitHub, often with GitHub Projects, organizes development, tracks issues and features and keeps the history of every change.
The models
Everything connects to Hermes (plus whatever additional tool I might use on a given day). Until recently, my Hermes sessions ran a powerful model, this month it was GPT 5.6 with extra high thinking. But since the latest DeepSeek V4 Flash came out, I realized it’s powerful enough to run long-horizon coding tasks autonomously, orchestrate implementation sessions, validate, test and deploy. It does everything I was getting from GPT 5.6 in my stack, only about 80 times cheaper (!!!), which is honestly absurd. So for a bit over a week now I’ve been using DeepSeek models for everything, except vision, where I use MiniMax M3, since DeepSeek is not multimodal. As of today, August 14, while I’m writing this, GLM 5.3 just came out, a model that competes directly with Fable and GPT 5.6, again at a fraction of the cost, so I’ll definitely be adding it to my workflow too (probably alongside Kimi K3).
Beyond intelligence, DeepSeek has a throughput (tokens per second) almost 2x faster than GPT 5.6, so I’ve also noticed faster task execution, both in orchestration and implementation.
And of course, as I said, cost is probably the most striking part, the savings are simply ridiculous. That said, I don’t use the model APIs directly. For open-source models I use the OpenCode Go plan, highly recommended for getting access to all the open-source models in a single subscription, hosted in the US, EU and Singapore.
The orchestrator
From Hermes, all that knowledge base is used to analyze what’s on the table, design the architecture, plan the implementation, structure the work, create the tickets, validate, research (both in my own sources and on the web) and, as a final step, prepare the prompts, whether for another orchestrating session or for implementation sessions.
This last part matters and it’s the most significant change in my workflow in recent months: I no longer prompt coding agents. In the last few months, a lot of “agent harnesses” and skills (or prompting strategies, you could call them) have appeared, like ultrawork, drill-me, etc. At some point, when choosing which skills to use or how to structure prompting depends on a human, it gets tedious, slow and imprecise, because a successful session needs a prompt that is long, clear and tightly scoped.
That’s where the Hermes Agent magic comes in. After reviewing the information sources, it has enough context to discuss with me what problems to tackle, in what order and with what priority. The biggest value of my Hermes Agent is probably not the tool itself (Hermes) but the skills, the working methodologies and the SOUL.md. For coding tasks I always use around 5 skills, plus the additional ones the case requires (frontend, Rust, testing), between 5 and 10 total. Those 5 main skills hold the manual of how I work: how to orchestrate, how to test, how to do QA, and any other task that comes up during a work session.
Those skills weren’t created overnight. They’ve been through a long iteration process and keep evolving with methodologies like GEPA, although, again, that deserves an article of its own.
The implementation
Until very recently, Codex was my main implementation tool, with the GPT 5.6 Sol and Luna models, together with a plugin built on OmO called LazyCodex, which basically brings OmO’s methodology to Codex.
As for Claude Code, I’ve never used it. I always used Anthropic models from opencode and OmO, until Anthropic banned using its coding subscription in third-party tools.
So, what do I use today? oh-my-pi (OMP), a harness built on Pi, which is incredible: it doesn’t just increase task success rates, it also does it with lower token consumption, which brings the cost down too!
OMP also has a lot of interesting features that make it stand out from other harnesses: goal mode, LSP, subagents, an advisor model watching the session to avoid deviations, hashline, memory and much more.
Hermes launches these OMP sessions (headless) in what we call lanes, waves of work where several sessions run in parallel, which speeds up implementation a lot. I also use tools like codegraph that reduce token usage even further and increase success rates.
Still, looking at the recently released deepseek harness, I’ll definitely be trying it out in the coming days.
The conclusion
It’s hard to put a name on this system. In June, people started calling it “Loop Engineering”, but that seems to be fading or getting ambiguous, others would simply call it a harness. At the end of the day, the label doesn’t matter. This world changes and evolves so fast (literally in weeks) that labels can’t keep up.
No workflow is perfect, there’s always room for improvement and errors, I’m constantly finding things to change, it’s a system in constant evolution and progress.
Along with this, there’s a new trend that has gained popularity in recent months: self-improving agents. Part of Hermes’ popularity, especially at the beginning, comes from its ability to update its own skills. That’s great, but sometimes neither that nor memory is enough. That’s where methodologies like GEPA and paradigms like Recursive Language Models (RLM) come in, something very innovative that I’m sure we’ll be hearing more and more about in the coming months.
The closing
I’ll probably write another post about this in a few months (or weeks, who knows). I’m excited to think about how my workflow will have changed by then. Some things are predictable, many others are totally unexpected.
At the end of the day, in this world the important thing is to stay up to date, not be afraid to try new things and, above all, not get married to any tool and not be afraid to leave one behind if necessary.