explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • Quick answers
  • What Foresight actually does
  • Why EmbeddingGemma 2 is the interesting part
  • How it fits Google's local AI push
  • How it compares with other local meeting tools
  • What to verify before trusting it with real meetings
  • What developers can take from it
  • A sensible way to trial it
  • Why this matters
  • Related reading
← Back to blog

explainx / blog

Google AI Edge Foresight: A Local Mac Meeting Notes App Built on Gemma 4

Google, Gemma, On-Device AI, Productivity, Privacy

Google AI Edge Foresight is an experimental Mac app that expands shorthand into full meeting notes, fully on-device. What it does and what to check.

Oct 8, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Google AI Edge Foresight: A Local Mac Meeting Notes App Built on Gemma 4

Google has released Google AI Edge Foresight, an experimental Mac app that listens to a meeting, expands your rough shorthand into fuller notes, and lets you search past transcripts and files, with all inference running on your machine. Google says it works offline and has no cloud subscription cost. It is the first consumer-facing app built on EmbeddingGemma 2, the open multimodal embedding model Google DeepMind shipped two days ago.

The details come from Google's developer blog post on EmbeddingGemma 2 and Google AI Edge. This post separates what Google has stated from what you still need to verify yourself.

Quick answers

table · 2 cols
QuestionAnswer
What is it?An experimental macOS meeting companion from the Google AI Edge team
What does it do?Enriches your shorthand notes using the live conversation, and searches transcripts, notes, images and documents
Which models?EmbeddingGemma 2 (740M parameters) and Gemma 4 models, per Google
Does it need the internet?Google says no, it works completely offline
Does it cost a subscription?Google says there are no cloud subscription costs
Which meeting tools work?Any, because it reads system audio and the microphone
Is it production ready?Google calls it experimental

What Foresight actually does

Google's post describes three jobs. The first is enhanced note-taking: you type short fragments during a call and Foresight fills them out with details retrieved in real time from the conversation itself. If you jot "pricing, Q4, Dana" it can attach what was actually said about pricing.

The second is live question answering. Google's demo captions mention "questions detected and answered live by AI," meaning the app notices when someone asks a question in the meeting and drafts an answer from your indexed material.

The third is cross-modal retrieval. Because EmbeddingGemma 2 maps text, audio, images and video frames into a single vector space, Foresight can search images, documents, transcripts and notes with one natural-language query. Google says this runs against your personal knowledge library without leaving the device.

The capture method matters. Foresight integrates directly with system audio and the microphone rather than joining your call as a bot participant. That is why Google says it works "out-of-the-box with any meeting platform." It also means other participants never see a notetaker in the attendee list, which is both convenient and a consent question, covered below.

Why EmbeddingGemma 2 is the interesting part

Most local meeting tools chain several models: speech-to-text, a text embedder for search, maybe a captioner for screenshots. Google's pitch for EmbeddingGemma 2 is that a single compact model replaces that chain. According to the same developer post, the full multimodal model needs about 567MB of active RAM on a Pixel 11 Pro, and the text-only weights about 191MB. On a MacBook M5 Pro GPU, visual embeddings take as little as 37.3 ms per image (roughly 26.9 images per second), measured with a budget of 70 vision tokens per image.

Our EmbeddingGemma 2 explainer covers the model sizes, the Apache 2.0 license and the Matryoshka truncation options in detail. The short version for this story: Foresight is a reference implementation showing what a private, always-available retrieval layer feels like when the embedding model fits comfortably next to a chat model on a laptop.

Google also says EmbeddingGemma 2 can act as a zero-shot decision engine, matching inputs against label descriptions in milliseconds, and the same post shows a MediaPipe Decision task evaluating 500 options per turn in under 100 ms. That is a developer feature rather than something Foresight advertises, but it hints at where the on-device stack is heading.

How it fits Google's local AI push

Foresight did not appear from nowhere. Google has been steadily building out a local stack around Gemma 4:

  • Gemma 4 12B brought a multimodal model that fits laptops with 16GB of memory.
  • The July Gemma 4 updates improved tool calling and attention speed.
  • The Google AI Edge Gallery app, which now gains Instant Media Search and Video Moments Finder demos powered by EmbeddingGemma 2, is the playground for these models on Android and iOS.

Foresight is the Mac counterpart: where the Gallery shows what a model can do, Foresight wraps it into a task people already perform every day. Google says its inference engine, LiteRT, powers ML Kit, MediaPipe Tasks, the Gallery and Foresight, so the same optimizations carry across.

How it compares with other local meeting tools

Foresight lands in a space where open-source projects have been working for months. Meetily is a privacy-first meeting assistant with local transcription. FluidVoice handles on-device dictation on macOS, and Files.md takes the local-first, plain-markdown route to notes.

What differs is the vendor. Foresight is closed-source from Google, though the models underneath (the Gemma family and EmbeddingGemma 2) are open-weight. If you want to inspect or modify the pipeline, the open projects still win. If you want something polished from the team that trains the models, Foresight is the new option. Google has not said whether the app's source will be released, so do not assume it.

What to verify before trusting it with real meetings

Privacy claims are the vendor's. Google says processing is fully local and that sensitive information remains secure on your device. That is plausible for a local-inference design, but the launch post offers no third-party audit. A practical check on macOS: run Foresight with a network monitor such as Little Snitch or the built-in Activity Monitor network tab and watch for outbound connections during a test meeting. Check whether it phones home for model downloads, analytics or update checks, and confirm which of those you can disable.

Recording consent. Capturing system audio is technically simple and legally uneven. Some jurisdictions require every participant to consent to recording. A bot-free recorder removes the visible signal that a call is being captured, so tell participants yourself.

Accuracy of enhanced notes. Expanding shorthand with retrieved context can hallucinate links between your fragment and the wrong part of the conversation. Treat the output as a draft and spot-check names, figures and commitments, exactly as you would with a cloud notetaker.

Hardware. Google's post says the app is a Mac download, and other coverage reports it is optimized for Apple Silicon. The post does not publish minimum memory or disk requirements, and we have not benchmarked it ourselves, so expect to find out by trying. If you already run a Gemma 4 model alongside other apps, watch memory pressure during long calls.

Experimental status. Google labels it experimental. Features can change or vanish, and there is no stated support commitment.

What developers can take from it

Even if you never install Foresight, the launch post doubles as a recipe. The ingredients Google lists are:

  1. LiteRT for the runtime, with a single .litertlm file running on CPU and GPU across supported edge platforms.
  2. MediaPipe Tasks, which adds EmbeddingGemma 2 support to the Embedder and Semantic Retriever tasks and handles image resizing, tensor normalization and tokenization for you, returning 768-dimensional vectors that can be truncated to 128 to 512 dimensions.
  3. ML Kit on Android, coming in the next few weeks per Google, with NPU acceleration and automatic model updates.
  4. Pre-quantized bundles from the Hugging Face LiteRT Community, plus a Colab and GitHub guide.

A minimal local-search architecture, following Google's own description, is: embed every note, transcript chunk and image into one vector space, store vectors in a local database (the Gallery uses SQLite), and retrieve by cosine similarity as the user types. Add a small Gemma 4 model on top to rewrite or answer using the retrieved snippets. That is the same shape as Foresight, and the privacy story comes from never calling a hosted API.

A sensible way to trial it

Start with a low-stakes recording rather than a client call. Run a short internal meeting with Foresight capturing system audio, jot a few shorthand fragments as you go, and compare the enhanced notes against your own memory of what was said. Check three things: whether names and numbers survived intact, whether the expanded fragments point at the right moment in the conversation, and whether search over older notes and images returns what you expect when you type a loose query.

Keep the network monitor running throughout, and note memory pressure if a Gemma 4 model is also loaded. If the first two meetings hold up on accuracy and traffic, move to longer calls; if not, you have lost little. Pair the trial with one of the open alternatives above so you have a baseline to judge against.

Why this matters

For people who pay for cloud meeting assistants, a free local alternative from Google changes the default calculation, particularly for legal, medical, HR and finance conversations where uploading recordings is a compliance headache. For builders, it is more evidence that the useful part of a meeting assistant, retrieval over your own material, no longer needs a server. The unanswered questions are quality and trust: how good are Gemma 4-sized notes compared with frontier cloud models, and will Google keep the app maintained?

We will update this post once independent hands-on tests appear. If you try it, the checks worth publishing are an outbound traffic log, a side-by-side note quality comparison against a cloud tool, and memory use during a one-hour call.

Related reading

  • EmbeddingGemma 2: open multimodal embeddings in 740M parameters
  • Gemma 4 12B: multimodal local model for 16GB laptops
  • Gemma 4 updates: flash attention and tool calling
  • Meetily: privacy-first local meeting assistant
  • FluidVoice: open-source macOS dictation
  • Files.md: local-first note-taking

Official sources: Google developers blog on EmbeddingGemma 2 and AI Edge and Google AI Edge documentation.

Details reflect Google's announcement on October 8, 2026 and may change as the app updates.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Jun 15, 2026

Gemma 4 Powers Open Duck Mini: Meet Autumn, the On-Device AI Robot Duck

At Google I/O 2026, two tiny bipedal robot ducks showcased Gemma 4 E2B running fully on-device—one on a Raspberry Pi 5, one on a Jetson Orin Nano—using multimodal inputs to see, hear, and speak in real time.

Sep 23, 2026

djev-run: A One-Command Way to Deploy DiffusionGemma-Jev on Google Cloud Run

Deploying DiffusionGemma-Jev — the open-source Jev clone built on Google's diffusion Gemma model — used to mean provisioning your own GPU. A new project, djev-run, cuts that down to a single gcloud command that spins up a Jev API-compatible endpoint on Cloud Run, scaling to zero when idle. Here's what it actually does and what it costs to run.

Sep 6, 2026

In-Browser LLM Fine-Tuning: Why Training on WebGPU Is a Bigger Deal Than Inference

Reports surfaced of a developer fine-tuning a language model entirely inside the browser using WebGPU — no server, no cloud round-trip for the training step itself. explainx.ai breaks down why backpropagation in a browser sandbox is a meaningfully harder claim than the in-browser inference we already cover, what realistic scope looks like, and what to check before you believe the demo generalizes.