3つのAIエージェント、2つの国、そしてあまりに不均一なWorld Wide Web

Three AI agents, two countries, and one uneven world wide web

3つのAIエージェント、2つの国、そしてあまりに不均一なWorld Wide Web

World Bankの公共調達データベースを題材に、MetaのMuse、AnthropicのClaude Cowork、OpenAIのGPT 6.1 Solの3エージェントへ英語(米国)とFarsi(Iran)で同一タスクを実行させた比較検証。Iran側は138件のN/AフィールドのうちGPTとMuseが埋められたのはわずか21件、米国は130件中51〜64件。公式政府サイトからの引用比率は米国76〜89%に対しIranは11〜22%で、TelegramやMediumなど信頼性の低い情報源が混入した。登録・アップロード段階ではClaudeは拒否、GPTは人間に委ね、Museは利用規約を示さず[email protected]名義でアカウントを登録した。

For an evaluator outside an AI lab, without that access, it is nearly impossible to fully make sense of an agent’s behavior. And if outside evaluators can only see partial trajectories, and any conclusions they draw can be dismissed for lacking complete information, what is the value of independent evaluation?

この日のほかの記事

2026-10-02