In this special episode of Cloud Wars, Bob Evans speaks with Garrett Lord, co-founder and CEO of Handshake, about one of the biggest challenges facing enterprises today: turning the extraordinary promise of AI into measurable business outcomes. Lord explains why AI agents have delivered dramatic productivity gains in software engineering but have yet to achieve comparable results across many other knowledge-work domains.
Making AI Agents Work
- AI’s ROI Gap Persists: Enterprises are enthusiastic about AI and are spending heavily on tokens, but Lord says meaningful returns remain uneven. Software engineering has emerged as the standout, with AI agents producing what he describes as two- to threefold productivity improvements and increasingly executing lengthy assignments autonomously. The challenge is translating that success into disciplines such as finance, manufacturing, retail, oil and gas, and insurance. In these areas, agents can generate presentations, emails, and briefing documents, but they aren’t yet consistently transforming complex, long-running workflows.
- Evaluations Define What ‘Good’ Means: Lord repeatedly returns to evaluations as the foundation for enterprise AI. Because models are nondeterministic, companies can’t assume an agent will reliably produce the desired result simply because it performed well once. Instead, businesses need to codify what successful performance actually looks like across their specific workflows. Lord compares an evaluation to a product requirements document: it creates a measurable baseline against which an agent can continuously improve. Once organizations can evaluate performance, they can improve their agents and harnesses, decide which models are appropriate for particular jobs.
- Long-Horizon Work Is the Frontier: Generating an email or presentation is fundamentally different from completing a 20-hour investment-banking assignment. Lord describes knowledge work as a complex trajectory involving data rooms, Outlook, Slack, Excel, Bloomberg, FactSet, colleagues, and potentially hundreds of individual actions. Humans continuously manage context while navigating those systems, but today’s AI agents still struggle to execute these long-horizon workflows reliably outside software engineering. That distinction helps explain why impressive AI demonstrations haven’t always translated into enterprise transformation.
The Big Quote: “The point that most enterprises are at right now is they want to bring agents into production beyond software engineering.”
More from Garrett Lord and Handshake:
Follow Garrett on LinkedIn and learn more about Handshake’s approach to enterprise AI at the following links: Handshake Enterprise AI, The Right Model for the Right Work, and Handshake AI.



