Here's how I'd carve it up — the real tastemakers are the places where a new model gets tried, ranked, and talked about within 48 hours of release.
Distribution layers that create instant adoption
- OpenRouter — probably the single best early signal. It's the default router for indie devs and apps, and its public token-share rankings by model and by app are the closest thing to a Billboard chart for LLMs.
- Hugging Face — for open weights, the trending page and download counts are the canonical proxy for mindshare. Also where fine-tunes and quants appear, which is what actually spreads a model.
- Cursor, Claude Code, Cline/Roo, Windsurf, Copilot — coding agents are the highest-value workload right now, so getting into their model picker (and their default) is a kingmaker moment.
Judgment layers that shape reputation
- LMArena — human preference elo; heavily gamed and criticized for style bias, but still the number every launch post cites.
- Artificial Analysis — the intelligence-vs-price-vs-speed charts that procurement people screenshot.
- Benchmark suites with cultural weight — SWE-bench Verified, Aider polyglot, ARC-AGI, GPQA. Less "platform," more shared scoreboard.
Discourse layers
- AI X/Twitter — a few hundred accounts effectively set the narrative in the first 24 hours; vibe checks here outrun benchmarks.
- r/LocalLLaMA — the open-weights kingmaker; a model that fails here doesn't get quantized, and unquantized means unadopted.
- Hacker News, and increasingly YouTube/Discord for the practitioner layer.
Lagging but decisive
- Bedrock, Vertex, Azure AI Foundry — they don't set trends, they ratify them. Availability there is the signal that enterprise budget will follow.
If I had to pick three to watch: OpenRouter usage share for real behavior, r/LocalLLaMA or AI Twitter for sentiment, and coding-agent default status for where the money is.