Ecosystem

Perplexity lets a model watch its own production

2 min read AI-generated

In OpenAI's new customer story, Perplexity cofounder Johnny Ho says Astra edits and monitors their live systems - and that he checks in on it far less often than he used to.

Featured image for "Perplexity lets a model watch its own production"

OpenAI published a new customer story, and it is more interesting than the format suggests. The voice in it is Johnny Ho, cofounder and Chief Strategy Officer at Perplexity.

The core, in his words: «We can have the model craft communications, edit real-world systems, and monitor our production software in a way that previous generations were not able to.»

Three things in one sentence: write communications, change live systems, watch production software. Two and three are where you stop reading and go back a line.

The model writes its own test doubles

Ho gets most concrete about testing. There is never time to test by hand, so he has Astra build a small testing program around an application. The model produces realistic responses, the kind another service would send back, say an LLM API or a connector. It stands in for the other end and then checks how the application behaves, all the way through the workflow.

Anyone who has written mocks for an API that does not exist yet knows why this is more than a party trick. This is exactly the work everyone postpones.

The second quote is the real one: «We’re actually able to trust it with full end-to-end systems and check in on it much less frequently than previous generations of models.»

One aside I liked: Ho notes that every time the model gets better at writing code, Perplexity’s search engine gets better too, because it can then write better programs that search the web and internal sources and summarize them tightly.

Why this matters to me

It is a vendor’s customer story, so take it with the appropriate salt. Still, here is someone saying out loud that he has dialed back supervision. Not «we prototype with it» but production systems and fewer check-ins.

That is the other half of this week. Amodei warns about agent swarms, Washington argues about liability, and meanwhile the line where an ordinary engineering team stops watching the model quietly moves. Both happened on the same day. The second one just happened with less noise.

Sources:

OpenAIAgentsEngineeringPerplexity