The Harness Effect - How Orchestration Design Sets the Token Economics of Enterprise Agentic AI
my team's work on WRITER Agent's harness
“Token-maxxing” was popular in March 2026 as a form of conspicuous consumption; it was assumed that the work you completed was linearly proportional to the amount of LLM tokens you used over a period of time. After about a month, everyone came to their senses once the bill got sent in and realized this was not the case. How should we move forward, not just as developers but as engineers building AI products useful to normal, cost-conscious people and businesses?
This was one of the principles behind WRITER Agent’s new harness; it has to deliver results for sales and marketing teams faster and at lower cost when using a frontier model, and at similar quality to a frontier model when using a smaller model. We achieved both and measured the results.