- Model Highlights: Muse-Glimmer-30B
A distillation of Muse Spark and optimized for local deployment
- Model Highlights: MiniMax H3
The omni-modal video & audio generator
- Why is your model being careful with email bodies but reckless with bank accounts?
Beyond Aggregate Risk: Role-Stratified Conformal Risk Control for LLM Tool Calls
- Model Highlight: Laguna S 2.1
A new heavyweight for agentic coding?
- Why force models to compute knowledge when they could just look it up?
Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
- Why build a bigger model when you can just loop twice for twice the power?
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
- Can an AI agent run the entire scientific method without human supervision?
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement
- Why do your coding agents keep getting lost in large repositories?
SWE-Explore: Benchmarking How Coding Agents Explore Repositories
- Are you still manually fighting with LaTeX and TikZ to create publication-quality figures?
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs
- Can your AI agent actually learn from its mistakes or just keep repeating them?
SkillOpt: Executive Strategy for Self-Evolving Agent Skills










