Very cool to see Pi build a durable agent harness too. I've been building in this space for quite some time myself [0][1] and it is a super interesting place of innovation. Less hype-y that on-your-machine coding agents, but all major players are building products in this space: LangChain Deep Agents, Vercel Eve, OpenAI Agents API, Anthropic Managed Agents, etc.
The main reasons are:
1) they are "durable", i.e. easier to make long-running in an unattended way, and easier to implement recovery, monitoring, etc
2) separating the harness from the compute brings safety and scaling benefits
It's quite an interesting place to hack on, because it's both a well understood problem but with so many wrinkles to it. I can't count how many earlier designs we chewed through before we ended up with the final one and I would not be surprised if we learn even more about it.
Yes, from your release post I can tell that you put a lot of thought into it!
I like your structured concurrency approach with tasks, which is similar to how I do it in Lightspeed too.
Also, the durable state implementation as documents is elegant! Question, though: why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?
Things are still not settled, and while the API looks like store i/o it's actually more similar to Immer's drafts, just with a different encoding, as JSON patch can't deal with the kinds of data we encounter in our workloads, at least not in a way that keeps memory and perf within some bounds.
The good thing is that this more low level API can be easily papered over with a nice sugary thing.
> why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?
We tried so many things. At one point it pulls in so much more complexity. At one point we had half of automerge's proxy system in there. In the end we felt like this is a reasonable line to draw, but we will see!
Pi is anyway a better option than all the others, because of it's focus on agnosticism. Vercel's SDK will work slightly better with it's AI gateway, OpenAi's Agent will work better with codex models, and so on.
Pi-durable makes pi a good acquisition target for Cloudflare - nothing like durable objects (with containers no less) really exists in other clouds. Wonder what the_mitsuhiko thinks about this.
One decision here which seems like a large break from the original pi is that Durable doesn't support branching conversation trees, it only supports conversation forks with ancestry information. Can anyone speculate (or confirm, if you happen to be Armin or Mario) why this is, and if that is necessary for the durable guarantees? The branching conversations are still an immutable data structure, so I can't see why this would be necessary, but perhaps I'm missing something.
I think it’s just for consistency. A fork and a tree navigation are the same operation conceptually. Now, unlike older Pi, a fork is not a copy of the session, it just has a pointer to the older session. The only losses that I can see are: now /resume shows every conversation rewind; and /tree is harder to implement. Neither of those is provided by Durable, so the gap is left to the implementor.
It's an interesting concept. This is half way to replicating pieces of Gastown. I like the idea, but I'm disappointed these tools still fail to address sandboxing as a first class citizen. I want to be able to declaratively set rules for what sandboxes agents execute in and mark context as tainted when untrusted etc. So far I still don't see any of these harnesses properly addressing this space. I'd be interested in knowing if it can be done through the extensibility of Pi, but since it operates directly on the trust layer, it feels like the type of thing that really needs native support.
Not familiar with the details of Pi Durable, but I have tinkered a bit with different sandboxing strategies for Pi. IMHO it would be hard to trust a sandboxing layer built into a harness that is so focused on being fully pluggable/moddable/self-improvable.
When I am using Pi to write extensions for Pi, I feel better running Pi wrapped in a separate os-level sandbox. I guess Pi could do it all, but I am content with how it is.
if you only have one level of trust then running the harness itself in a sandbox and leaving it at that is fine. This works for coding. For more complex enterprise style scenarios it stops working. Say you have an agent reading emails for you to action high priority ones. You have to assume it is going to get prompt injected constantly. But you want to have an escalation pathway for a high priority email, so somewhere you need a tool that can modify state in a database. You can't give that trust to the email reading one. So you need a higher level agent that can spin up a low trust sub-agent, get an output from it, and then feed the sanitised output into a different agent that has rights to update the database. This is obviously simplified / toy scenario, but it just illustrates that there are different trust levels, and different agents need to be authorised to do different things.
I don’t really understand the infinitely running case, but I do have schedule agents to do things like open PRs when some events happen, or triage alerts every day
Something I've had Chat GPT do for a while was writing "Daily Presidential Briefings" for me. Basically non-clickbait, well-summarized, priority-ordered news.
one example: I have several cron jobs that monitor different projects and notify me via my claw agent in Telegram. Whenever it sends me a message, I can ask it to address the issue within the same conversation.
This stuff is really complicated. Just trying to build a harness coordinating multiple instances of vanilla pi has been a bit of a nightmare. I'm not sure if the huge added complexity is worth it but kudos for trying and labelling as experimental.
super interesting. I have so many half-considered questions... like sandboxing (it seems like it is BYO). would love to have some kind of policy engine.. perhaps an integration with https://github.com/NVIDIA/openshell in the form of an extension?
also, I see most of the durability promise comes from persisting JSON documents locally and minimizing the amount of context/data kept in-memory, even during SQLite mode. while this makes sense, my own experiments with a process that relied on a JSONL-based event store have led me to prefer keeping things in-memory to avoid all the friction with I/O.. am I crazy for preferring just a straight .db file being persisted?
Brilliant. I wish the the durable application state wasn't restricted to just the json documents though. There should be some sort of integrated way of implementing the outbox pattern so external stores can be synchronized with the conversation state.
The commit on the root conversation picks out the last agent answer id from the transcript, and durably schedules a task that then syncs it to postgres. inside the task, you fetch the answer by id and send it over to postgres indempotently.
What's missing here is sugar, basically a hook that runs inside each commit so the outbox write is atomic with the state change, with ordered delivery, and possibly a durable change feed with cursors.
I like the multi-user bit the most. Should make it easier to build my remote control tool, as something I've had to hack around is not being able to use ACP while the TUI is active in an instance.
Cross-post from the 1.0 thread. I’m currently building a harness for Slack to support our on-call and support channels. It’s been working great so far.
The harness is built on top of the Pi SDK. I initially used Codex, but Pi seems more hackable, and I like that it’s vendor-agnostic by default.
Running it on Kubernetes works, but dealing with the JSONL session files and making sure sessions survive pod interruptions adds some complexity. I’m using DBOS for that right now, which works well, although it still feels like overkill.
This came at just the right time. I’m looking forward to removing the pieces I no longer need and simplifying the architecture. Thanks Pi team!
Very cool to see Pi build a durable agent harness too. I've been building in this space for quite some time myself [0][1] and it is a super interesting place of innovation. Less hype-y that on-your-machine coding agents, but all major players are building products in this space: LangChain Deep Agents, Vercel Eve, OpenAI Agents API, Anthropic Managed Agents, etc.
The main reasons are:
1) they are "durable", i.e. easier to make long-running in an unattended way, and easier to implement recovery, monitoring, etc
2) separating the harness from the compute brings safety and scaling benefits
3) easier to make multi-player.
[0] https://github.com/smartcomputer-ai/lightspeed
[1] https://github.com/smartcomputer-ai/agent-os/
It's quite an interesting place to hack on, because it's both a well understood problem but with so many wrinkles to it. I can't count how many earlier designs we chewed through before we ended up with the final one and I would not be surprised if we learn even more about it.
Yes, from your release post I can tell that you put a lot of thought into it!
I like your structured concurrency approach with tasks, which is similar to how I do it in Lightspeed too.
Also, the durable state implementation as documents is elegant! Question, though: why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?
Things are still not settled, and while the API looks like store i/o it's actually more similar to Immer's drafts, just with a different encoding, as JSON patch can't deal with the kinds of data we encounter in our workloads, at least not in a way that keeps memory and perf within some bounds.
The good thing is that this more low level API can be easily papered over with a nice sugary thing.
> why directly write/read to the store, why not abstract it and do more of a reducer/redux pattern and hide the persistence of the documents?
We tried so many things. At one point it pulls in so much more complexity. At one point we had half of automerge's proxy system in there. In the end we felt like this is a reasonable line to draw, but we will see!
I'm loving it. Already converted a few of my smaller tools to it, and am ripping out the guts of piclaw to replace them (in time)
Pi is anyway a better option than all the others, because of it's focus on agnosticism. Vercel's SDK will work slightly better with it's AI gateway, OpenAi's Agent will work better with codex models, and so on.
Pi-durable makes pi a good acquisition target for Cloudflare - nothing like durable objects (with containers no less) really exists in other clouds. Wonder what the_mitsuhiko thinks about this.
One decision here which seems like a large break from the original pi is that Durable doesn't support branching conversation trees, it only supports conversation forks with ancestry information. Can anyone speculate (or confirm, if you happen to be Armin or Mario) why this is, and if that is necessary for the durable guarantees? The branching conversations are still an immutable data structure, so I can't see why this would be necessary, but perhaps I'm missing something.
I think it’s just for consistency. A fork and a tree navigation are the same operation conceptually. Now, unlike older Pi, a fork is not a copy of the session, it just has a pointer to the older session. The only losses that I can see are: now /resume shows every conversation rewind; and /tree is harder to implement. Neither of those is provided by Durable, so the gap is left to the implementor.
A fork is branch right? This is how pi's branches are built.
> The entire source code, without tests, is about 15,000 lines, which comes out to about 150,000 tokens with GPT and about 250,000 with Claude.
Woah, that big of a difference when it comes to token counting?
Some of the delta is newish:
https://openrouter.ai/blog/insights/opus-47-tokenizer-analys...
Yeah don't get anyone going on the labs different tokenization schemes lol
It's an interesting concept. This is half way to replicating pieces of Gastown. I like the idea, but I'm disappointed these tools still fail to address sandboxing as a first class citizen. I want to be able to declaratively set rules for what sandboxes agents execute in and mark context as tainted when untrusted etc. So far I still don't see any of these harnesses properly addressing this space. I'd be interested in knowing if it can be done through the extensibility of Pi, but since it operates directly on the trust layer, it feels like the type of thing that really needs native support.
Not familiar with the details of Pi Durable, but I have tinkered a bit with different sandboxing strategies for Pi. IMHO it would be hard to trust a sandboxing layer built into a harness that is so focused on being fully pluggable/moddable/self-improvable.
When I am using Pi to write extensions for Pi, I feel better running Pi wrapped in a separate os-level sandbox. I guess Pi could do it all, but I am content with how it is.
if you only have one level of trust then running the harness itself in a sandbox and leaving it at that is fine. This works for coding. For more complex enterprise style scenarios it stops working. Say you have an agent reading emails for you to action high priority ones. You have to assume it is going to get prompt injected constantly. But you want to have an escalation pathway for a high priority email, so somewhere you need a tool that can modify state in a database. You can't give that trust to the email reading one. So you need a higher level agent that can spin up a low trust sub-agent, get an output from it, and then feed the sanitised output into a different agent that has rights to update the database. This is obviously simplified / toy scenario, but it just illustrates that there are different trust levels, and different agents need to be authorised to do different things.
After looking at so many options, that is also my take.
These should be decoupled.
Maybe I need nono in one context and smolvm in another or both.
I would not want to trust the harness to self policy.
I agree in using a separate OS-level sandbox or a VM. Better to have the option for modularity.
However, for ease of use, it is nice for harnesses to by default run with sane and safe sandboxing setup. Then give the option to disable them.
What are people using these infinitely-running agents for?
The extremely basic use case is that I'm doing something at work and it's not finished yet when I leave the office.
Or my laptop crashes, ugh.
Yes, if I could ssh into a random server it'd be fine. But I can't.
I don’t really understand the infinitely running case, but I do have schedule agents to do things like open PRs when some events happen, or triage alerts every day
Something I've had Chat GPT do for a while was writing "Daily Presidential Briefings" for me. Basically non-clickbait, well-summarized, priority-ordered news.
one example: I have several cron jobs that monitor different projects and notify me via my claw agent in Telegram. Whenever it sends me a message, I can ask it to address the issue within the same conversation.
This stuff is really complicated. Just trying to build a harness coordinating multiple instances of vanilla pi has been a bit of a nightmare. I'm not sure if the huge added complexity is worth it but kudos for trying and labelling as experimental.
super interesting. I have so many half-considered questions... like sandboxing (it seems like it is BYO). would love to have some kind of policy engine.. perhaps an integration with https://github.com/NVIDIA/openshell in the form of an extension?
also, I see most of the durability promise comes from persisting JSON documents locally and minimizing the amount of context/data kept in-memory, even during SQLite mode. while this makes sense, my own experiments with a process that relied on a JSONL-based event store have led me to prefer keeping things in-memory to avoid all the friction with I/O.. am I crazy for preferring just a straight .db file being persisted?
Brilliant. I wish the the durable application state wasn't restricted to just the json documents though. There should be some sort of integrated way of implementing the outbox pattern so external stores can be synchronized with the conversation state.
You can already (sort of, kind of) do that via a task (please excuse the agent slop, it's midnight and it's been a long day):
``` const SyncToPostgres = defineTask<{ entryId: string }, { phase: "send" }, void>({ kind: "app.sync-postgres", version: 1, initial: () => ({ phase: "send" }), phases: { send: async (task, runtime, context) => { const entry = await runtime.read(/* the entry */); await postgres.upsert("messages", { id: task.input.entryId, ...entry }); // idempotent by id await runtime.commit(() => ({ status: "terminal", outcome: { status: "completed" } }), context); }, }, });
```The commit on the root conversation picks out the last agent answer id from the transcript, and durably schedules a task that then syncs it to postgres. inside the task, you fetch the answer by id and send it over to postgres indempotently.
What's missing here is sugar, basically a hook that runs inside each commit so the outbox write is atomic with the state change, with ordered delivery, and possibly a durable change feed with cursors.
Thanks for the input!
Nice! I made my own version of this for Pi but I’m excited to see if I can just replace it lol
I like the multi-user bit the most. Should make it easier to build my remote control tool, as something I've had to hack around is not being able to use ACP while the TUI is active in an instance.
> A requestId makes a submission exactly-once, so a client that retries after a crash gets the original submission back instead of asking twice.
Sounds like a bug under a false assumption. Just having an ID cannot alone guarantee exactly once semantics AFAIK.
Explain this
in Effect, please
Cross-post from the 1.0 thread. I’m currently building a harness for Slack to support our on-call and support channels. It’s been working great so far.
The harness is built on top of the Pi SDK. I initially used Codex, but Pi seems more hackable, and I like that it’s vendor-agnostic by default.
Running it on Kubernetes works, but dealing with the JSONL session files and making sure sessions survive pod interruptions adds some complexity. I’m using DBOS for that right now, which works well, although it still feels like overkill.
This came at just the right time. I’m looking forward to removing the pieces I no longer need and simplifying the architecture. Thanks Pi team!
I’m shoving this into Agent Substrate on Kubernetes with a different storage interface
That code font is painful to read, no syntax highlighting and extremely pixelated. It's retro but an eyesore.