Interesting. When it comes to auditing, I think there is also another layer to be verified - the intent. We have tooling for static analysis and automated tests, but drift in intent or vision is hard to detect.
After leaving agents unsupervised for longer periods of time, the biggest questions tend to be "how did it arrive at this conclusion?" and "did we drift from the initial vision?". Typically these questions are hard to answer as you only review the final state of the workflow.
I recently built an issue tracker that does offer time-traveling as a core feature. Somewhat unexpectedly this turned out to provide a much sought for overview allowing for a better workflow audit, as detailed here:
https://dev.to/ljtn/vision-drift-addressing-the-next-problem...
Would be interested to hear your thoughts. I guess vision drift would classify as something that is hard to reverse, because you often don't realize it has happened until much later.
There are so many other factors. Will people be pissed if they expected your output and got AI slop instead? Will you lose context on something that you really need to understand?
Interesting. When it comes to auditing, I think there is also another layer to be verified - the intent. We have tooling for static analysis and automated tests, but drift in intent or vision is hard to detect.
After leaving agents unsupervised for longer periods of time, the biggest questions tend to be "how did it arrive at this conclusion?" and "did we drift from the initial vision?". Typically these questions are hard to answer as you only review the final state of the workflow.
I recently built an issue tracker that does offer time-traveling as a core feature. Somewhat unexpectedly this turned out to provide a much sought for overview allowing for a better workflow audit, as detailed here: https://dev.to/ljtn/vision-drift-addressing-the-next-problem...
Would be interested to hear your thoughts. I guess vision drift would classify as something that is hard to reverse, because you often don't realize it has happened until much later.
I work on my MacBook in direct terminal sessions or in a container with no safeguards.
I save the container for well-planned units of work that can easily be repeated and validated on a branch.
I use direct sessions when I expect and the plan dictates that Claude might go off target and I need to hold the steering wheel.
Fable has increased my use of the container option.
How do you draw the line between a human review at level 1 and a human safety gate at level 2?
There are so many other factors. Will people be pissed if they expected your output and got AI slop instead? Will you lose context on something that you really need to understand?