Six announcements in fourteen months, each advancing software operation — and each stopping at a different boundary.
With fallback computer-use tiers. The typed tier is opt-in, app by app.
Agents run in a separate, contained workspace — away from the desktop the person is using.
In preview, with governance at the session level.
Built with OpenAI and NVIDIA. Separate from the desktop.
Typed App Intents remain vendor-declared — and it is Siri alone, on Apple’s operating systems alone.
Governance at the session level.
Every platform validated the gesture. None combined installed-base reach, shared-desktop control, every desktop OS, and a check at the exact target. That combination is what InvokeFlow delivers, today. We extend these platforms — we don’t compete with them.
Agents can reason well enough to do real work. They still can’t do it where the work is. In a bank, an insurer, a hospital or on a sales floor, the work lives in installed applications — and the agent has no sanctioned way in. Not because it can’t operate them. Because no one can bound what it does, or prove what it did.
Two clocks converged. Agent capability crossed the threshold where enterprises want to deploy. And the vendors shipping agents carved out the sectors where the highest-value work lives. The demand exists; the supply is prohibited from serving it. Meanwhile the installed base isn’t getting rewritten — it’s in maintenance, source-lost, regulator-validated, or simply working.
Screen-driven automation answers a different question — how do we mimic a hand on a screen? — and inherits every weakness of that answer: brittle when the screen changes, a seat instead of a grant, a log instead of a record, and often a machine that isn’t the user’s. APIs cover the cooperating minority. Neither was built for an actor a compliance officer has to sign for.
Software invocation. Call what installed software can already do, in place, per command, with a record. Adoption is what people do. Invocation is what software does.
— David Richard and Raj Madugula, co-founders
A peer-reviewed systems study ran the same AI model on the same real desktop tasks twice. In one run the agent had to find and click. In the other it could name what it wanted and let a layer handle the mechanics.
Task success, before and after the agent could invoke instead of click.
Fewer interaction steps for the same work.
Of successful tasks finished in a single model call.
The AI didn’t get smarter. It got a better way to reach the software. Source: Wang, Li & Chen, EuroSys ’26.
Steps that never needed judgement run deterministically, with no model call. The token toll on those is zero.
Invoking a command is one step; finding and clicking it is several. Every step you don’t take is one you don’t pay for.
Most tasks finish in a single model call instead of a look-click-look cycle. Faster for the person waiting, cheaper for the budget.
Point InvokeFlow at the applications it lives in. It runs governed and on the record from the first command.