AntLingAGI, the model team at Ant Group, has made Ling-3.1-flash available free on OpenCode. It is an open-weights model with 560 billion total parameters and 25 billion active per token, and it ranks second among open-weights assistants on Design Arena's Mobile App Arena at 1,207 Elo.
Free access to a model near the open-weights frontier is a genuine opportunity for anyone running a coding agent. It also comes with a few gaps — a context-window discrepancy, an unconfirmed license, and benchmark claims that are mostly secondhand — worth knowing before you commit a workflow to it.
TL;DR — what people are asking
| Question | Answer |
|---|---|
| What is it? | Ling-3.1-flash, an open-weights model from AntLingAGI (Ant Group) |
| Size? | 560B total parameters, 25B active per token |
| Where? | Free on OpenCode |
| Context window? | 262K on OpenCode; developer previously cited up to 1M |
| Mobile App Arena? | No. 2 open-weights, 1,207 Elo; No. 17 overall |
| Other scores? | GDPVal-AA v2.1 1,673 Elo; FrontierSWE 75.16; HealthBench Professional 65.35 |
| Optimized for? | Healthcare, coding and general productivity |
| License? | Described as open-weights; exact terms unconfirmed |
| Catch? | Free tiers change; benchmark framing is secondhand |
What it is
The numbers describe a large mixture-of-experts design. A model with 560B parameters but only 25B active per token routes each token through a small subset of its experts. That is why a model this large can be hosted cheaply enough to offer free: compute per token tracks the 25B active figure, not the 560B total.
For readers who want the background on sparse models and why active parameters drive cost, our Kimi K3 architecture write-up and open-weights guide explain the trade-offs.
AntLingAGI is the model effort of Ant Group, the company behind Alipay. Some coverage describes it as an Alibaba affiliate and a sibling to the Qwen family; the safest statement is that it is an Ant Group effort in the same broader ecosystem. If you track that ecosystem, our Qwen coverage is the best context.
The scoreboard
| Measure | Reported result |
|---|---|
| Mobile App Arena (Design Arena) | No. 2 open-weights assistant, 1,207 Elo; No. 17 overall |
| GDPVal-AA v2.1 | 1,673 Elo |
| FrontierSWE | 75.16 |
| HealthBench Professional | 65.35 |
Three notes on reading these.
- Arena rankings are preference-based. Mobile App Arena measures how people judge generated mobile app output, so it tells you about a specific kind of work — building mobile UIs — and is not a general intelligence score.
- The "close to GPT-5.6 Sol and Opus 5" line is a characterization. It comes from coverage of the launch, not from a head-to-head I ran. Treat it as a reason to test, not a conclusion.
- No. 17 overall is the more sober number. Second among open models is a real result; seventeenth among all models is a reminder that the closed frontier is still ahead. Epoch's index puts the top of the field at Opus 5.5 with a score of 167, and open models sit well below it.
For a primer on how to weigh scores like these, see how to read AI benchmarks.
The context-window gap
OpenCode lists a 262K context window for Ling-3.1-flash. The developer has previously described capability up to 1 million tokens. Both can be true — a model can be trained for longer context than a host chooses to serve — but the discrepancy matters if your workflow depends on long inputs.
Practical guidance: design for 262K. If you need to feed an entire large repository in one shot, test the limit yourself with a realistic input, and treat any longer window as unverified until the host or developer confirms it.
Why "free on OpenCode" matters
OpenCode is an open-source coding agent, and hosting a strong open model there lowers the cost of running agent loops. Coding agents burn tokens quickly, so a free capable model changes the economics of experimentation. Our OpenCode guide walks through setup, and running open-source models locally with OpenCode covers the self-hosted route if you want to avoid a hosted free tier altogether.
Free-tier caveats worth checking
- Rate limits. Free tiers throttle. A long agent run can stall midway.
- Data handling. Free hosted models may log prompts or use them differently from paid plans. Do not paste proprietary code until you have read the terms.
- Durability. Free offers end. Keep your agent configuration portable so you can swap models.
- Provider routing. Hosted models can be served from regions or providers you do not control.
How to try it this week
- Install OpenCode and select Ling-3.1-flash from the model list, following the OpenCode setup guide.
- Run a bounded task from your own backlog: a refactor, a test-writing pass, a small feature. Keep it short enough to fit comfortably inside 262K.
- Compare against your current default on the same task. Judge on correctness, number of turns, and diff quality.
- Test a mobile UI task if that is your work, since the arena result is specific to it.
- Skip sensitive repos until you have read the data terms.
- Record the result. A five-line note per task beats a general impression a week later.
If you want more context on how open and closed models trade off in practice, read the Mozilla analysis of the open-weight frontier gap.
A fair test plan for a free model
Free models are tempting to adopt on enthusiasm. A short, structured trial keeps the decision honest.
| Step | What you do | What you record |
|---|---|---|
| Pick tasks | Choose five tasks from your recent work: a bug fix, a test-writing job, a refactor, a UI component, a documentation edit | Task description and acceptance check |
| Run both models | Run Ling-3.1-flash and your current default on each task with the same prompt | Turns taken, tool-call errors, whether tests pass |
| Blind review | Have someone, or you after a delay, rate the diffs without knowing which model wrote them | Score from one to five |
| Check long context | Feed one task near the 200K range | Whether quality holds or degrades |
| Check failure modes | Note any invented APIs, ignored instructions, or truncated outputs | A list of repeat offenders |
| Decide | Adopt for the tasks where it matched, keep the default elsewhere | A routing rule you can write in one sentence |
The last row is the useful outcome. Most teams end up with a routing rule — use the free model for drafts and bulk edits, use a stronger model for design decisions and difficult debugging — rather than a wholesale switch. That rule captures most of the savings while keeping quality where it matters.
Reading the arena result sensibly
An arena score such as 1,207 Elo is a preference measure over many head-to-head comparisons. Two practical consequences follow. A few dozen Elo points is a small difference, so being second rather than third among open models should not drive a decision. And an arena built around mobile app generation rewards visual polish and completeness, which are not the only things that matter in a codebase you have to maintain.
What people are asking
Is Ling-3.1-flash better than Qwen?
The sources do not compare them directly. They are from related ecosystems, and your own task test will tell you more than a ranking.
Can I run it locally?
At 560B total parameters, local deployment needs serious hardware even if only 25B are active, because the full weights must be available. Most people will use the hosted version. Check the model card for weights and quantizations.
Is it good for coding?
Coverage lists coding among its optimized areas and reports a 75.16 score on FrontierSWE. That is promising and unverified by me; test on your own repos.
Should I worry that it is from a Chinese company?
Some teams have policies about model provenance and data residency. If yours does, check those policies before routing code through a hosted free tier. See our coverage of censorship and provenance concerns in popular Chinese open models for one lens on the issue.
Why is the arena rank different from the overall rank?
Mobile App Arena is a specialized arena where it ranks higher among open models; the overall figure compares it with every model on the leaderboard, including closed frontier systems.
Honest limitations
- I relied on launch coverage and did not benchmark the model myself.
- The license and exact commercial terms are unconfirmed.
- The 262K versus 1M context discrepancy is unresolved.
- Free access terms can change without notice.
Bottom line
Ling-3.1-flash is a large, sparse, open-weights model that is free to try on OpenCode and ranks second among open models on a mobile-app arena. That makes it a low-cost addition to your evaluation set. Verify the license, plan for 262K of context, and keep sensitive code off free tiers.
Related on explainx.ai
- OpenCode: open-source AI coding agent guide
- Run open-source models locally with OpenCode
- Choose open-weight vs closed AI models
- Mozilla: the open-weight frontier gap
- Qwen's 3 billion downloads
- Kimi K3 open weights
- How to read AI benchmarks
- Claude Opus 5.5 tops the Epoch Capabilities Index
Specs and rankings reflect launch coverage as of October 3, 2026. Verify the model card and OpenCode listing before relying on them.
