explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: what was announced
  • What is Evidence Finder?
  • Where does Namazu fit in the pipeline?
  • What does 96.4% on the physician exam mean?
  • What is Namazu, and how is it adapted to Japanese?
  • Why does this matter beyond Japan?
  • What are the limits and open questions?
  • How to evaluate a similar tool yourself
  • What people are asking
  • Related reading
← Back to blog

explainx / blog

Sakana Namazu Powers Aillis Evidence Finder: A Japanese LLM Scores 96.4% on the Physician Exam

Sakana AI, Namazu, Medical AI, Japan, Sovereign AI

Part of Open-Weight Models

Sakana AI says its Japanese LLM Namazu now powers Aillis's Evidence Finder for doctors and scored 96.4% on Japan's 120th physician exam. What it shows and what it does not.

Oct 9, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Sakana Namazu Powers Aillis Evidence Finder: A Japanese LLM Scores 96.4% on the Physician Exam

Sakana AI announced on October 9, 2026 that its Japanese-focused language model, Sakana Namazu, has been adopted by Evidence Finder. Evidence Finder is a literature search tool for physicians from Aillis Inc. (アイリス株式会社). The announcement carries one headline number: a Namazu-based evaluation model scored 96.4% on Japan's 120th national physician examination.

The post is a customer story, not a new model release. It is useful for two reasons. It shows a Japanese-tuned model working as the answer layer in a real medical-search product. And it shows, by the company's own wording, why an exam score alone should not be read as proof that an AI is ready for clinical work.

An earlier request to explain "Namazu / AILLIS" confused the names. AILLIS is the customer, Aillis Inc. Namazu is Sakana's model.

TL;DR: what was announced

table · 2 cols
QuestionAnswer
Who is involved?Sakana AI (model: Namazu) and Aillis Inc. (product: Evidence Finder)
What happened?Evidence Finder adopted Sakana Namazu to write comparative, cited answers for doctors
Headline number96.4% on Japan's 120th national physician exam (February 2026), per Aillis
Caveat from SakanaAn exam score shows one facet of knowledge and reasoning, not clinical usefulness
What does the model do?Compares and synthesizes papers that Aillis's algorithms select, then writes the answer
What checks the sources?A Verify function that confirms cited papers exist
Is there a new model?No. The post describes the current Namazu
Pricing or API details?None in the post
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

What is Evidence Finder?

Evidence Finder is a service from Aillis, a Japanese medical AI company. According to Sakana's post, it "searches literature from databases such as PubMed in response to a doctor's question and generates answers while naming the sources." It has a Verify function that automatically checks whether the papers cited by the AI actually exist. The design goal is to let doctors "reach the primary literature quickly."

The need is plain. Sakana's post says clinicians must find, read and compare findings from papers and guidelines that update every day, but doctors whose main job is treating patients have limited time for literature research. A tool that finds papers and writes a cited summary saves time, but only if the citations are real. Hallucinated references are the known failure of general chatbots. The Verify step targets exactly that. For background on why models invent sources and how to catch them, see explainx.ai's guide on why AI models hallucinate.

Japanese trade press reported earlier that Aillis agreed to use Sakana's model in Evidence Finder on September 14, 2026, according to a Yakuji Nippo summary. Sakana's own post is dated October 9.

Where does Namazu fit in the pipeline?

The post splits the work between two parties.

  1. Aillis algorithms choose relevant papers for the doctor's question. This is the search and verification layer that Aillis built.
  2. Sakana Namazu compares and integrates those papers and writes the answer.

This is a retrieval-augmented design. The retrieval layer finds the evidence. The language model writes and organizes it. If you build similar systems, the split is worth copying. It limits what the model can claim to what the retrieved papers say. explainx.ai covers the tradeoffs in RAG versus agentic RAG and in grounding with RAG or fine-tuning.

Sakana says the adoption reflects its Japanese answer ability and the medical performance described next. In Sakana's words, the model "naturally" fits the workflow of getting doctors to trusted primary literature.

What does 96.4% on the physician exam mean?

Stethoscope tied by a thread to a booklet, standing for Sakana Namazu answering physician questions with cited literature

The post says the Namazu-based evaluation model for Evidence Finder scored a 96.4% correct rate on the 120th National Medical Licensing Examination, held in February 2026. A Japanese trade summary of Aillis's own announcement gives the raw score as 482 out of 500 points.

Aillis says this is the highest score among domestic foundation models designated by the Ministry of Economy, Trade and Industry's GENIAC program, among those with publicly confirmable results. The footnote explains the method: as of September 1, 2026, Aillis compared the physician exam results from the 118th exam onward that are available in each company's published materials. That is a comparison by Aillis, among published results, and not an independent leaderboard.

Sakana adds its own warning. The post says: "The national exam result shows only one facet of knowledge and reasoning ability, and does not directly show usefulness in clinical practice." The decision to adopt Namazu also rested on a qualitative review of real answers, which found practical usefulness.

That is the right way to read it. A licensing exam tests recall and reasoning on multiple-choice questions. Real clinical questions are open, messy and tied to a patient. The exam score is a screening signal. It is not an outcome study.

What is Namazu, and how is it adapted to Japanese?

Sakana describes Namazu as a Japanese-specification LLM. It starts from a strong open model. The post names the problem that often follows when you train a high-performing open model further on Japanese: "catastrophic forgetting," where the model loses existing knowledge and reasoning, or becomes over-optimized for narrow tasks.

Sakana says Namazu addresses this with its own post-training technique. It keeps the base model's reasoning, knowledge and coding ability on major benchmarks while strengthening adaptation to Japanese language, Japanese cultural and business context, and neutral responses. The current Namazu maintains the base model's performance in mathematical reasoning, general knowledge and reasoning, and coding. It beats the base model on Japanese instruction following, translation and evaluations about Japan-specific context.

The post states the goal: not "just a model that is good at Japanese," but a way to make "the capabilities of the world's best open models usable in Japan." Sakana also says it is building a base where inference completes inside Japan, including data handling and operations, so Japanese companies and organizations can choose it more easily.

The post does not name the base model or give benchmark tables. explainx.ai covered the earlier product surfaces in Sakana Chat adds Fugu, new Namazu and code execution. That August post noted that no new public benchmark numbers came with the update. The 96.4% exam score is the first concrete domain number attached to a Namazu deployment that we have seen.

Why does this matter beyond Japan?

There are three lessons for people who build with AI.

Domain wrappers beat raw models. The value in Evidence Finder is the pipeline: selected papers, a writer model, a source check. Namazu is a part of it. A smaller or regional model, well tuned for language and tone, can be the right writer when the retrieval is strong.

Language adaptation without losing skill is a real engineering problem. Catastrophic forgetting is why many local-language fine-tunes feel weaker at reasoning. Sakana claims its post-training avoids that. The claim is not independently benchmarked in this post, but the target is correct, and the approach connects to the guide on what fine-tuning is and when it helps.

Sovereignty has a medical use. Japan wants models that run domestically with clear data handling. Healthcare is a place where that preference is strong. This links to the broader debate in What Is Sovereign AI? and to Europe's parallel path with models such as Aleph Alpha Kolibri.

Sakana is also widening where Namazu applies. The post lists human resources, marketing, customer support and internal research as areas that need Japanese systems and business-practice knowledge. It says Namazu is strongest at "accurately understanding Japanese instructions and materials, conducting consistent work from information retrieval through organization and explanation." Interested companies can consult via a form.

What are the limits and open questions?

  • The score is vendor-reported. Aillis ran the comparison. Sakana repeats it with attribution. No independent test is cited.
  • It is an evaluation model. The post says "the evaluation model of Evidence Finder using Sakana Namazu" scored 96.4%. The production model may differ. The post does not say.
  • No clinical outcome data. There is no study of whether doctors using the tool make better decisions or save a measured amount of time.
  • No error analysis. The post does not report how often Verify catches a fake citation, or how often the summary misreads a paper.
  • Base model is unnamed. You cannot reproduce or compare the adaptation without it.
  • Medical advice caution. Evidence Finder is described for physicians. It is not a patient-facing diagnosis tool. People should not use any AI search summary in place of medical advice. For a patient-side view of AI and health records, see ChatGPT Health and your medical records.
  • Source language. We read the Japanese original of Sakana's post.

How to evaluate a similar tool yourself

Two check stamps beside a rubber stamp, showing how to verify Evidence Finder style citations one by one

If you are choosing or building an evidence-search assistant, five checks give you more than any exam score.

  1. Citation existence. Take 50 answers and check each citation by hand. Count fakes.
  2. Citation support. For real citations, check that the paper says what the answer claims.
  3. Recency. Ask about a guideline updated in the last six months. See if the tool finds the new version.
  4. Refusal behavior. Ask something with no good evidence. A good tool says so.
  5. Language quality. Ask in the user's real language, with the real jargon, and have an expert read the output.

What people are asking

Is Namazu a medical model? No. Namazu is a general Japanese-specification model. Evidence Finder is the medical product. The exam score belongs to the Namazu-based evaluation model.

Does Namazu replace doctors' search? It is designed to shorten the path to primary papers, with the doctor still reading the source.

Is this the same as Sakana Fugu? No. Fugu is Sakana's multi-model orchestrator, covered in Sakana Fugu: one API to orchestrate the others. Namazu is the Japanese-first model.

Related reading

  • Sakana Chat adds Fugu, new Namazu and in-chat code execution
  • Sakana Fugu: one model API to orchestrate all the others
  • Sakana Fugu Max and Ultra v2
  • What Is Sovereign AI?
  • Why AI models hallucinate and how to catch it
  • RAG versus agentic RAG
  • What is fine-tuning?

Primary sources: Sakana AI: Sakana Namazu adopted by Evidence Finder (Japanese); Aillis Inc.; Yakuji Nippo summary of the September 14 agreement (Japanese).


Figures come from Sakana AI's October 9, 2026 post and Aillis's published comparison as of September 1, 2026. They are company claims and may be updated.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Aug 14, 2026

Sakana Chat Adds Fugu, New Namazu, and In-Chat Code Execution

On August 13, 2026 Sakana AI put Fugu and a new Namazu generation into Sakana Chat, then added the piece that actually changes the product: sandboxed Python, a side-panel for HTML/Word/slides, and image plus document attachments. This is Japanese-first vibe coding in a browser — not a new frontier API SKU.

Oct 9, 2026

Ecosia Drops Mistral for Open-Weight Models: What Politico Reported and What It Means

Politico reported on October 6, 2026 that Ecosia is dropping Mistral as its AI model supplier and moving to open-weight models, some of them Chinese, hosted by Melious. The headline many people saw was a new Ecosia and Mistral partnership. The reporting says the opposite. This post separates the facts from the framing.

Oct 6, 2026

Mistral Large 4 "Le Chonk": 1T Parameters, 49B Active, Open Weights Due October 27

Mistral released a preview of Large 4, nicknamed Le Chonk, on October 6, 2026: a natively multimodal mixture-of-experts model with about 1 trillion parameters and 49 billion active per token. The API is live today, the open weights are promised for October 27, and the lab claims the strongest open-weight results from the US or Europe. Here is what the numbers say, what is still unverified, and what to do before the weights land.