Back to perspectives
PERSPECTIVES
Voice, Code, and Autonomy: Gradient Descending in September
September Gradient Descending sessions took us from voice AI in Barcelona to Google’s codebase in Munich and back to London for autonomous software engineering.
Oct 7, 2026
7 Min Read
Ecosystem Insights

Share
There was a moment at our Factory AI session in London when the conversation turned to a software engineering task that had been running for 22 days. Not an agent spinning for an afternoon while someone occasionally checked whether it had broken something, but a system keeping track of the original objective, breaking the work down, executing it and checking its own progress over more than three weeks. A few days earlier in Munich, we had been having almost the opposite conversation with Stoyan Nikolov from Google: what happens when AI makes it incredibly easy to create more code inside organisations that already have an enormous amount of it?
Somewhere between those two discussions is a pretty good picture of September at Gradient Descending. We came back from the summer with sessions across London, Munich and Barcelona, covering voice AI, huge codebases, inference and autonomous software engineering.We wanted the subjects to span very different areas . Gradient Descending was never supposed to be a conference chopped into smaller pieces; the fun of it is putting twenty or thirty people around a table, choosing one technical problem and seeing where the conversation goes. This month, it went to some particularly interesting places.
Barcelona: voice AI is getting very good but production is still complicated.
We started in Barcelona with Sergio Espeja from HeyDiga and Ismael Ordaz from SLNG, just as the latest generation of real-time speech models was making the old cascade-versus-speech-to-speech debate much more interesting. The quality of speech-to-speech has moved quickly when conversations feel more continuous, interruptions are more natural and some of the tricks traditionally used to make cascaded systems feel human are becoming unnecessary. It is easy to listen to one of these systems and assume that the architecture question has effectively been answered.
Once you move from a great demo into production, there are plenty of reasons to want the control that comes with a cascaded architecture. Teams can choose the speech-to-text model, the LLM and the voice independently, inspect each component when something goes wrong and make different decisions around deployment, cost and data. Evaluation is also much easier when the system can be pulled apart: if an agent gives a bad answer, you can work out whether it heard the user incorrectly, reasoned badly or produced the wrong output. With end-to-end audio models, those boundaries become much harder to see.
The conversation eventually moved away from architecture and towards expectations - people are surprisingly tolerant of other people making mistakes. A human on the phone can misunderstand us, ask us to repeat ourselves or give us the wrong information and, within reason, we accept it. When an AI does the same thing, we tend to see it as a failure of the product. That makes reliability, traceability and control particularly important in voice, especially once these systems start handling conversations where getting something wrong actually matters.

Munich: what if the problem becomes too much code
A week later, we were sitting around a table in Munich with Google’s Stoyan Nikolov. Google is an interesting place to discuss coding agents because its definition of a large codebase is slightly different from most companies’: billions of lines of code, huge dependency graphs and years of engineering infrastructure designed to keep all of it functioning.
The obvious promise of coding agents is productivity. Engineers can produce more, migrations can happen faster and increasingly substantial pieces of work can be handed to agents. But at Google’s scale that creates a slightly uncomfortable question, what happens if we become dramatically better at producing code without becoming equally good at maintaining it? More code means more changes to review, more dependencies to understand and more software that somebody, human or agent, needs to look after in five or ten years. The bottleneck doesn’t necessarily disappear but it moves somewhere else.
That took the discussion away from the coding agent itself and towards the environment it works in. A model operating inside a huge repository needs more than access to source files. It needs to build graphs, tests, dependencies, ownership information and all the engineering structure that tells an experienced developer how a codebase actually works. As agents take on more maintenance and migration work, there is an interesting possibility that repositories themselves will start changing: not simply code written for machines to execute and humans to understand, but codebases structured so that agents can navigate and maintain them too.

London: how much autonomy can you actually give an agent
We ended September with Ryan Lieber and Dmytro Yaroshenko from Factory AI and a discussion that pushed the coding-agent idea considerably further. Factory’s ambition is to bring autonomy to software engineering and the distinction between autonomy and productivity came up repeatedly. The aim isn’t simply to give every developer a better copilot, but to build systems capable of taking responsibility for progressively larger parts of the software development lifecycle.
Factory has been testing that idea internally. Their system can take in signals from customer feedback, telemetry, ideas and other parts of the company, identify work that needs doing, execute it and validate the result. Around 12% of changes can now be approved and merged automatically when seven months earlier, every PR still received a human rubber stamp. What was particularly interesting, though, was how little of the discussion was about finding a smarter model.
Dmytro described this as a progression from prompt engineering, to context engineering, to environment engineering. An LLM can decide that it has completed a task and produce a very convincing explanation of why everything works, a deterministic test is considerably less persuadable. Factory currently runs 82 deterministic checks against its codebase, creating the kind of “back pressure” that forces an agent to confront mistakes rather than simply conclude that the job is finished. For longer-running work, responsibility is separated across agents that orchestrate, execute and validate, with state kept outside the model so the system can continue returning to the original objective rather than relying on an ever-growing context window.
That is what made the 22-day mission interesting. The headline is obviously that an AI system can keep working towards the same software engineering objective for more than three weeks. The more important part is everything Factory had to build before that was useful - checks, validation, external memory, orchestration and clear conditions for what “finished” actually means. As agents become capable of working for longer periods, giving them more time is relatively easy. Giving them enough structure that you can trust what they do with that time is much harder.
The models will keep changing, probably faster than most of us can plan around. What feels more durable are the engineering questions underneath them: how you evaluate them, what context you give them, how you constrain them, what infrastructure they run on and how you know when they have actually done what you asked. Those are the conversations we want Gradient Descending to keep making room for, ideally with twenty people around a table, a few disagreements, and considerably more questions at the end than we had at the beginning.

PERSPECTIVES
Related articles
Keep up with Earlybird and our portfolio companies.




