explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR: the questions people are asking
  • What LittleBit is
  • How it works
  • What the numbers say, and what they do not
  • What the repository lets you do today
  • LittleBit-2: what changed
  • How it compares with other small-model routes
  • What this means for what you build or pay
  • Limits and open questions
  • Summary
  • Related reading
← Back to blog

explainx / blog

Samsung LittleBit: A 13B LLM Under 1 GB, and What the Paper Really Shows

Local AI, Quantization, Samsung, Open Source AI, AI Research

Part of Local AI

Samsung LittleBit squeezes Llama2-13B to under 0.9 GB at 0.1 bits per weight. What the NeurIPS paper shows, the CC BY-NC license, and what you cannot run yet.

Oct 9, 2026·11 min read·Yash Thakker
add explainx.ai
go deep
Samsung LittleBit: A 13B LLM Under 1 GB, and What the Paper Really Shows

A headline on October 8, 2026 said Samsung had shrunk a 13B language model to under 1 GB. The claim is real. It comes from a NeurIPS 2025 paper and a code release from Samsung Research. But the story is easy to misread. This is a compression method that needs a training run. It is not a model you can download and chat with.

This post covers what LittleBit does, how the numbers in the paper should be read, what the repository lets you do today, and how it compares with other small-model routes on explainx.ai. We read the paper abstract and the repository README. We did not run the training code, so every result below is the authors' own.

Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.

TL;DR: the questions people are asking

table · 2 cols
QuestionShort answer
What shipped?Official code for LittleBit (NeurIPS 2025) and LittleBit-2 (ICML 2026) in SamsungLabs/LittleBit
Is the 13B in 1 GB claim true?The paper abstract says Llama2-13B compresses to under 0.9 GB at 0.1 bits per weight
Can I download a model?No ready checkpoint is listed in the README. You train your own
What does it cost to use?A QAT run. The README example uses 5 epochs on C4 and WikiText data
Can I use it commercially?No. The repo is CC BY-NC 4.0
Is the 11.6x speedup measured?The paper calls it a potential speedup relative to FP16
Which models work?OPT, Llama and Llama 2/3, Phi-4, Qwen2.5 and QwQ, Gemma 2 and 3, Qwen3

What LittleBit is

LittleBit is a framework for what the authors call sub-1-bit compression. Normal quantization stores each weight in 4, 3 or 2 bits. Even the 1-bit models in the news store about one bit per weight. LittleBit stores less than one bit per weight on average.

The README gives the one-line version: it "compresses large language models into the sub-1-bit regime by factorizing each dense weight matrix into low-rank latent factors, binarizing those factors, and restoring magnitude information through lightweight learned scales." The authors report support down to 0.1 bits per weight, "while preserving the original model architecture at inference time."

The first paper, "LittleBit: Ultra Low-Bit Quantization via Latent Factorization," is by Banseok Lee, Dongkyu Kim, Youngcheon You and Youngmin Kim. It was first posted on arXiv on May 30, 2025, and the current version is v5 from February 5, 2026. So the method is not new. What changed in October 2026 is the attention: the repository now carries the follow-up, and social posts picked it up.

How it works

Three ideas carry the method.

1. Low-rank factorization, then binarization

A weight matrix W is approximated as the product of two thin matrices U and V. In a normal low-rank approximation, U and V stay in 16-bit floats. LittleBit keeps latent full-precision copies during training, but at inference time it uses only their signs, +1 or -1. A matrix of signs needs one bit per entry. Because U and V are thin, the total count of entries is far below the count in W. That is how the average drops under one bit per original weight.

2. Multi-scale compensation

Signs lose magnitude. LittleBit adds learned FP16 scale vectors along three axes: rows, columns and the latent dimension. The paper calls this multi-scale compensation. These scales are small compared with the matrices, and they let the binary factors approximate the original values.

3. Dual-SVID initialization and residual compensation

Training from a bad starting point is unstable at this precision. The paper introduces Dual Sign-Value-Independent Decomposition (Dual-SVID). It starts from a truncated SVD of the real weights and maps the signs and magnitudes into the learnable parameters. A second, residual branch then models what the first branch missed. You can see both branches in the figure below.

LittleBit factorized linear layer with primary and residual binary branches for sub-1-bit LLM compression

Figure: the LittleBit linear layer, with a primary and a residual branch of binary factors and learned scales. Source: the LittleBit paper (arXiv 2506.13771), Samsung Research. Reproduced for commentary with credit to the authors.

Training then runs as quantization-aware training (QAT). The model sees the binarized forward pass during training and learns to compensate. The README names SmoothSign as the quantization function and LittleBitLinear as the module.

The XOR claim

Some posts say LittleBit "replaces floating-point multiplication with bitwise XOR." That is a fair description of why binary factors are cheap. When both operands are signs, a multiply becomes an XOR plus a popcount. It is a hardware argument, not something the repository ships as a kernel. Keep that distinction in mind when you read the speedup figure.

What the numbers say, and what they do not

Two bars comparing bit widths, showing sub-1-bit LLM quantization using far fewer bits per weight

The paper reports these results:

table · 3 cols
ClaimSource and wordingCaveat
About 31x memory reduction at 0.1 BPWAbstractRelative to FP16
Llama2-13B under 0.9 GBAbstractSee the arithmetic note below
0.1 BPW beats leading methods at 0.7 BPW on Llama2-7BAbstractCompared with prior sub-1-bit work
32B model at 0.3 BPW stays strongIntroductionCompared against STBLLM, the prior leading sub-1-bit technique
Potential 11.6x inference speedup vs FP16AbstractThe word is potential, not measured end to end
Robust below 0.5 BPWFigure 1 captionPrior method degrades sharply there

The comparison baseline is STBLLM, which the paper calls the leading sub-1-bit technique. The paper says LittleBit at 0.1 BPW on Llama2-13B is competitive with, and surpasses, STBLLM at 0.55 BPW. The tested range is 1.3B to 32B parameters.

An arithmetic note

A 13B model at exactly 0.1 bit per weight would need about 0.16 GB for those weights. FP16 is 26 GB, and 26 divided by 31 is about 0.84 GB. So the under-0.9 GB figure is consistent with the 31x reduction, but not with 0.1 bits across every parameter. Our reading is that some parts of the model stay at higher precision, for example embeddings, the output head or the scale vectors. The abstract does not give the breakdown. If your budget depends on the exact size, read the paper's memory table before you plan.

Quality is not full quality

Sub-1-bit does not mean lossless. Perplexity and zero-shot scores at 0.1 BPW sit below the FP16 model. The paper's claim is that LittleBit loses less than prior methods at the same size. For a chat assistant or a coding agent, expect a weaker model than the original. A 13B model at 0.1 bits is not a 13B model at 16 bits.

What the repository lets you do today

The repository SamsungLabs/LittleBit had 148 stars, 17 forks and 5 commits when we read it. It is a research codebase.

table · 2 cols
ItemDetail
LicenseCC BY-NC 4.0
Python3.12 recommended
PyTorch2.8.0 with CUDA 12.4 in the README
TransformersPin 4.51.x to reproduce paper results
Trainingpython -m main with QAT, single GPU or DeepSpeed ZeRO-3
Evaluationeval.py with WikiText2 and C4 perplexity, plus BoolQ, PIQA, HellaSwag, WinoGrande, ARC and OpenBookQA
Supported modelsOPT, Llama and Llama 2/3, Phi-4, Qwen2.5, QwQ, Gemma 2 and 3, Qwen3

A minimal training command from the README looks like this:

bash
CUDA_VISIBLE_DEVICES=0 python -m main \
  --model_id meta-llama/Llama-2-7b-hf \
  --dataset c4_wiki \
  --save_dir ./outputs/Llama-2-7b-LittleBit-2 \
  --num_train_epochs 5.0 \
  --quant_func SmoothSign \
  --quant_mod LittleBitLinear \
  --residual True \
  --eff_bit 1.0

The --eff_bit flag sets the effective bits per weight. The README says the code is designed for 1.0 down to 0.1. To turn on LittleBit-2, add --use_itq True.

There is no llama.cpp, MLX or Ollama export in the README. That matters for the "run locally" angle. You can evaluate a checkpoint with the repository's eval.py, but a tight, fast phone runtime for these factorized layers is not part of the release. For ready local runtimes, see what llama.cpp is and how to run GGUF models.

LittleBit-2: what changed

The follow-up, "LittleBit-2: Maximizing the Spectral Energy Gain in Sub-1-Bit LLMs via Latent Geometry Alignment," is by Banseok Lee and Youngmin Kim and appeared at ICML 2026. It targets the initialization stage. It applies Internal Latent Rotation with Joint Iterative Quantization (Joint-ITQ). The goal is to align the SVD-derived latent factors with the binary hypercube before QAT starts.

The README states two useful properties. First, it "modifies initialization only; the deployed factorized layer remains the same," so there is no extra inference cost. Second, it is opt-in, so you can compare runs with and without it.

How it compares with other small-model routes

A large block simplified into a small tiled block, the idea behind running a 13B LLM in under 1 GB

LittleBit is one of several routes to small local models. They trade effort for size in different ways.

table · 4 cols
RouteBits per weightNeeds training?Ready to run today?
Post-training 4-bit (GPTQ, AWQ, GGUF Q4)about 4NoYes
Ternary and 1-bit modelsabout 1 to 1.58Trained from scratch or convertedSome are, for example PrismML Bonsai 27B
Ternary packing tricksabout 1.48VariesResearch, see BITCOS
1-bit GGUF of a huge modelabout 1NoYes, see Kimi K3 1-bit GGUF
LittleBit0.1 to 1.0Yes, QATCode only

If you want a smaller model this week, a 4-bit GGUF or an existing 1-bit release is the practical path. The complete guide to AI model quantization explains those formats. Ternary work such as Neutrino-1 8B sits near the 1-bit line that LittleBit goes beneath.

What this means for what you build or pay

Most readers will not run LittleBit. The result still changes how to think about small models.

  1. Memory is not the only floor. The size wall for local models moves down as methods improve. A 13B-class model in a budget of under 1 GB is no longer absurd on paper. That matters for phones, wearables and embedded boards.
  2. QAT is the price. Training-free formats win on convenience. LittleBit wins on size. Pick based on whether you can afford a GPU training run.
  3. Watch the license. CC BY-NC 4.0 means research and personal use only. A company cannot ship a product on this code without a separate agreement with Samsung.
  4. Wait for kernels. The speedup depends on fast binary kernels. Until a runtime ships them, the 11.6x number stays theoretical.

Limits and open questions

  • Quality at 0.1 BPW. The paper compares against prior sub-1-bit methods, not against FP16 on real tasks like coding or tool use.
  • No public checkpoints in the README. We found none listed. Check the repository for updates.
  • Training cost. Five epochs of QAT on a 7B or 13B model needs real GPU time. The README does not state hours.
  • Reproduction pins. The authors tell you to use transformers 4.51.x. Newer releases may change results.
  • Independent replication. We found no third-party replication yet. This is an authors-only result for now.
  • Hacker News. A search for the topic on Hacker News on October 9 returned no thread about LittleBit, so we have no community reports to add.

Summary

Samsung's LittleBit is a credible research method for pushing LLMs below one bit per weight. The under-1 GB claim for Llama2-13B comes from the paper itself. The release is a training codebase under a non-commercial license, not a finished local model. Use it to learn where compression is heading, and keep using 4-bit or 1-bit GGUF releases for work you need to ship.

Related reading

  • What is AI model quantization? A complete guide
  • What is llama.cpp? Install and run GGUF models
  • PrismML Bonsai 27B: 1-bit Qwen3.6 on iPhone
  • BITCOS: ternary LLMs below 1.58 bits
  • Kimi K3 1-bit GGUF on a Mac Studio
  • Neutrino-1 8B ternary weights

Primary sources: the LittleBit repository and the LittleBit paper on arXiv. The HuggingNews summary of the story is here.

Specs, license terms, and repository statistics are accurate as of October 9, 2026. Check the repository for new checkpoints and kernels before you plan around this method.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 3, 2026

llama.cpp Adds Decision Model Support for Five Open Models

On October 2, 2026 the ggml team announced decision-model support in llama-server. You POST state plus typed questions to /v1/systemone and get option probabilities in one forward pass. Five open models ship as GGUF today — and the local runner map just got a third path next to Ollaya and Python/MLX stacks.

Sep 29, 2026

Jeff: Home-Trained Jev-Compatible 0.8B Decision Models

A Hacker News front-page project from firelex ships Qwen3.5 and Gemma fine-tunes that speak Jev's request format locally. The 2B's five-benchmark panel score is 83.1% against Jev's published 83.0% on a different sample; the 0.8B is ~28 ms on an M4 Max. This post is about Jeff specifically — not Kev, OpenJev, or Laya — and what the thread actually argued.

Aug 27, 2026

GLM-5.3 Open Weights Delayed — Z.ai Misses Its Own Aug 28 Target

Z.ai promised GLM-5.3's open weights roughly two weeks after its August 14 launch — and its own Hugging Face placeholder page counted down to August 28. That date passed without a release. Here's what was actually promised, what shipped instead (GLM-5.3-Flash, which reportedly topped OpenRouter), and what the slip means if you're planning around self-hosting GLM-5.3.