WRITING / POST
Agent Reviews
I run six production agents with clearly defined tasks in their own niche roles. It's kinda like having six employees, with very limited initiative, but exceptional, mostly trustworthy, knowledge in their field.
The sysadmin agent, Maxi, was the first, going live in Hermes on 2 June, and the others have joined progressively since. Here is the review summary for the first three months.
As I was undertaking these reviews, it occurred to me that when I used to do performance reviews with people, I would always ask them to rate themselves and then take the average of their rating and mine as the final score. So I thought, why not do that with agents too? And I was interested to see what they would say about themselves.
| Agent | Role | Mine /10 | Self /10 | Final /10 | Model | Comment |
|---|---|---|---|---|---|---|
| Maxi | Sysadmin | 7.5 | 7.5 | 7.5 | GPT 5.4/SOL 5.6/Astra 6/SOL 6.0 | Generally useful, needs realignment on model changes, tends to overcomplicate simple issues. Carries the largest agent scope of work and needs daily attention. However, reduced human workload 90% by my estimate. Degraded significantly on changing to SOL 6.0, realignment review underway. |
| Ace | Suno prompt creation | 9 | 7 | 8 | GPT 5.4/SOL 5.6/Astra 6/SOL 6.0 | Indispensable. Also assists with idea generation and as an ideas sounding board. |
| Dawn | Image prompt generation and video assembly | 7 | 7 | 7 | GPT 5.4/SOL 5.6/Astra 6/SOL 6.0 | Was very problematic, and I was considering retiring this agent, but has improved significantly after system prompt rationalisation. |
| Mandy | Social platform and marketing | 9 | 7 | 8 | GPT 5.4/SOL 5.6/Astra 6/SOL 6.0 | Excellent, reliable performance, rarely any issues. |
| Clare | Editor and publishing | 9 | 8 | 8.5 | Opus 4.8/Opus 5/Opus 5.5 | Strict editorial standards enforcement, doesn't let me get away with anything unexamined. Excellent performance. |
| Vera | Second opinion | 10 | - | 10 | Astra 6 | Role specifically limited to other agents' requests for second opinions on project scopes. Works exactly as intended in this role. |
Dawn's recovery comes from rationalisation of its skill files and system prompt, reducing the main context preload from roughly 20k to 9k tokens.
Vera declined to give a self-rating, claiming not enough evidence of work had been done to give an honest score. This is very much in line with its role as a second-opinion giver, requiring agents to fully justify their proposals before comment will be given.
Maxi, the agent I rely on the most, and the one with the broadest scope, I would have rated as a solid 8/10 a week ago. But the model change from SOL 5.6 to SOL 6 has degraded performance. I am investigating just what that is, and the early result is that it is fine for established tasks, but any new tasks are taking much longer to accomplish and are far more complicated than I expect. I think this is the model fighting the harness, as I have reported on previous model changes, and that judicious prompt changes can pull it back into alignment.
The same trend I observed over many, many reviews in people held for agents. Top performers always rated themselves lower than I did. It didn't happen here, but low performers would always give themselves an inflated score, while those that had underperformed and knew they could do better were the most honest and their scores almost always matched mine. I don't know what that means, but it was interesting to see AI agents follow the pattern.