explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

On this page

  • TL;DR
  • Where Maigret sits in the toolchain
  • Features that matter operationally
  • Deployment patterns
  • Ethics, law, and enterprise reality
  • Start with a question you are authorized to answer
  • Distinguish an account match from an identity match
  • Interpret failed checks as failed checks
  • Keep automated summaries attached to evidence
  • Build a report another analyst can reproduce
  • Pilot the workflow before adding it to recurring jobs
  • Related on explainx.ai
  • Sources
← Back to blog

explainx / blog

Maigret: open-source username OSINT across 3,000+ sites (soxoj/maigret)

OSINT, Maigret, Python, Cybersecurity, Open source, Infosec

Maigret builds a dossier from a single username—async checks across thousands of sites, HTML/PDF/graph reports, web UI, Tor/I2P—MIT-licensed Python 3.10+ with an auto-updating site database.

May 6, 2026·9 min read·Yash Thakker
add explainx.ai
go deep
Maigret: open-source username OSINT across 3,000+ sites (soxoj/maigret)

Maigret is a Python-first OSINT utility: feed it a username, and it fans out across a maintained catalog of social and niche sites, collecting public profile signals and packaging them into reports you can hand to an analyst or a ticket. The canonical repo is soxoj/maigret (~23k GitHub stars at the time of writing—refresh the page; the number moves).

This post is a capability overview for blue teams, researchers, and engineers who already think in terms of SOTL-style “find the same handle elsewhere” workflows—not a playbook for abuse.

TL;DR

table · 2 cols
QuestionAnswer
One-liner installPython 3.10+, then pip install maigret → maigret YOUR_USERNAME
ScaleREADME cites 3,000+ sites; default runs skew toward ~500 high-traffic entries unless you pass -a or --tags
OutputsHTML, PDF, XMind-style, JSON/NDJSON, CSV, TXT, --graph D3 graph
Web UImaigret --web <port> or docker run -p 5000:5000 soxoj/maigret:web
Stealth / regionTor, I2P, generic HTTP/SOCKS proxies
LicenseMIT
Weekly digest3.5k readers

Catch up on AI

Curated AI updates on agents, skills, and MCP — delivered to your inbox. Unsubscribe anytime.


Where Maigret sits in the toolchain

Maigret is complementary to manual triage and commercial OSINT suites (the upstream README’s Used by section names several vendors—verify current integrations on their sites). For teams, Maigret is the hackable variant you can script in CI or notebooks.

  • Breadth-first enumeration — quickly answer “does this handle exist on major and long-tail sites?”
  • Metadata harvest — pull bios, links, and cross-IDs that seed recursive searches (Maigret documents expansion from discovered usernames).
  • Reporting — one command can emit investigator-friendly HTML/PDF instead of a wall of logs.

Profile parsing is powered in part by socid_extractor for structured IDs from public pages.

Features that matter operationally

Site database hygiene. Maigret ships a bundled dataset and can fetch updates from GitHub (README: roughly once per 24 hours when online). Site-specific claimed vs unclaimed test pairs and --self-check exist so maintainers can tame false positives—a chronic issue in any username checker.

Tag and geography filters. --tags photo,dating or --tags us narrow the blast radius when you already know the persona’s likely vertical.

Permutation mode. --permute builds handle variants (e.g. from name parts), useful for typo-squat and alias hunting—also easy to misuse; pair with policy.

Parse a URL. --parse boots a search from an existing profile URL, handy when onboarding from a single IOC.

Deployment patterns

CLI on an analyst laptop — simplest path for IR retainers.

Docker — soxoj/maigret:latest for batch jobs; soxoj/maigret:web when you want a shared UI on a jump host (bind localhost or put it behind SSO; never expose raw OSINT UIs to the public internet without controls).

Embeddable library — README positions the CLI as a thin wrapper over async APIs you can import; see Read the Docs for library usage and options.

Ethics, law, and enterprise reality

Maigret is neutral tooling. The README’s disclaimer is explicit: lawful, educational use; you own GDPR/CCPA/employment obligations. For employee monitoring or vendor due diligence, run everything through legal and data-retention policy—public does not mean permissionless in every jurisdiction.

False positives still happen: shared usernames, bots, and homoglyphs can implicate the wrong person. Treat hits as leads, not verdicts.

Start with a question you are authorized to answer

For a defensive trial, use your own public handle or an organization-owned account you maintain. Write down the purpose before running the search: identifying forgotten company profiles, checking impersonation leads, or reviewing an account inventory. A clear purpose determines which results matter and prevents a broad report from becoming an unstructured collection of personal information.

A useful first question is "Which public profiles need manual review?" That is more defensible than "Who is this person?" The tool checks names across sites; identity is a separate inference. Keep the search task narrow enough that you can inspect the results and explain why each retained item belongs in the report.

The upstream repository and quickstart documentation describe the maintained commands and site database. Confirm the installed version and available options before copying an older command. Site counts and detection rules can change without changing the underlying question you are investigating.

Distinguish an account match from an identity match

Imagine that your company's public brand handle is northstar-demo. The tool finds a profile with that name on a forum and another on a photo site. The matching string establishes a lead. It does not establish that your organization controls either account, that the accounts belong to the same person, or that their profile statements are accurate.

Inspect the profile itself, its links, and the organization's known account records. An official website linking to a profile provides different evidence than a user-written biography claiming affiliation. Record the kind of evidence instead of flattening both into "confirmed." If a result remains ambiguous, preserve that ambiguity.

Common handles are especially easy to overinterpret. A short word, initials, or a popular nickname may be independently used by many people. Similar avatars can be copied. A profile may be abandoned, transferred, or impersonated. The analyst's responsibility is to determine what the available evidence actually supports.

Use separate columns for discovered URL, username match, source observations, affiliation evidence, and review status. This small structure makes it harder for a raw discovery to silently become an identity assertion when the report passes to another teammate.

Interpret failed checks as failed checks

A timeout or blocked request is not evidence that the account does not exist. A login wall, rate limit, changed page layout, or temporary service failure can all prevent a site check from reaching a meaningful result. Keep these outcomes separate from a verified negative response.

If many sites fail at once, inspect the local environment and network before drawing conclusions about the target handle. If one site produces suspiciously many matches, inspect its claimed and unclaimed behavior. A generic success page or custom error page can mislead an automated checker when the site changes.

For repeatable evaluation, maintain a small set of accounts you own and know exist, plus test names known to be unclaimed on the specific sites being checked. Revisit the labels when accounts are created or removed. The test is about the tool's detection behavior at a time and place, not a promise that the internet inventory is permanently complete.

Do not expand a blocked search into more intrusive access simply to make the report look comprehensive. Document the gap and choose an appropriate authorized source or manual review path. A useful report can include unresolved items when it explains why they remain unresolved.

Keep automated summaries attached to evidence

If you add an AI summarizer, provide the retained observations and explicit review statuses. Ask it to distinguish verified affiliation, probable association, and unresolved matches. The generated summary should point back to the underlying URLs and analyst notes so a reviewer can inspect its reasoning.

A model should not infer employment, location, or sensitive characteristics from a shared username alone. It can help organize a report, but it should not manufacture the missing identity link. Test the summary using an example with two unrelated accounts sharing the same handle; the expected result should preserve the distinction.

Check every consequential statement in the final summary against the retained evidence. If the model says the organization owns a profile that the analyst marked unresolved, revise the statement and the summarization instruction. Fluent wording can turn a tentative lead into a confident accusation unless the workflow catches that transformation.

For an agent integration, make the output schema reflect these boundaries. A field called "account found" is narrower than a field called "person identified." Clear field names reduce the chance that a downstream system treats uncertain discovery as proof.

Build a report another analyst can reproduce

Record the tool version, search date, scope, database state where available, and exact search input. Retain the result URLs and the observations necessary for the task. If a page later changes, the report should still explain what was observed and how the conclusion was reached.

Keep the report's audience in mind. An incident ticket may need only the profiles requiring review and the next action. A technical evaluation may need failures and false positives too. Avoid exporting a large dossier by default when a short account-inventory note answers the authorized question.

Define who may access the report and how long the work requires retaining it. Public source material can still become sensitive when aggregated. Follow the organization's applicable process for investigations and data handling rather than assuming the tool's software license establishes permission to collect or distribute every profile.

Pilot the workflow before adding it to recurring jobs

Run a small trial, inspect every result, and measure how much analyst work it creates. A broad search producing hundreds of low-confidence leads can be less useful than a focused check producing a few reviewable items. Choose automation around the task's actual need, not the largest advertised site count.

For recurring checks, define what constitutes a meaningful change and how someone reviews it. A new match should create a lead with context, not automatically label an account malicious or send accusations. Retain a path to correct mistaken associations and update the organization's known-account inventory.

The practical outcome is an evidence-backed review queue: what was found, what remains uncertain, and what an authorized owner should inspect next. Maigret accelerates discovery; the value of the workflow comes from careful interpretation and a report whose conclusions remain traceable to the sources.

Related on explainx.ai

  • Agent skills security — why automated tooling needs governance
  • What is MCP? — wiring structured tools into agent loops without leaking scope
  • Claude mythos preview — model-assisted defensive workflows in context

Sources

  • Repository: github.com/soxoj/maigret
  • Quick start: maigret.readthedocs.io — Quick start
  • Documentation: maigret.readthedocs.io
  • Site list: sites.md in repo
  • socid_extractor: github.com/soxoj/socid_extractor
  • PyPI: pypi.org/project/maigret
  • No-install option: Telegram bot (per README)
  • Maintainer commercial contact: maigret@soxoj.com (per README)

Site counts, Docker tags, and CLI flags change frequently. Treat this article as May 6, 2026 context—read maigret --help and the upstream CHANGELOG before production rollouts.

Spotted something out of date? Let us know.
Yash Thakker

Written by

Yash Thakker

Yash is an AI expert with over 300K learners. Join his workshops →

View Yash Thakker in People in AI →

Related posts

Oct 7, 2026

Anthropic Expands Its Cyber Verification Program Into Three Tiers

On October 6, 2026 Anthropic announced an expanded Cyber Verification Program with Defense, Red Team and Specialized tiers, and folded Project Glasswing into it. This post explains who qualifies, what each tier relaxes, what stays blocked, and what security teams should prepare before applying.

Oct 6, 2026

South Korea Bank Hacks: What AI Actually Did, and What Builders Should Learn

South Korea has launched a police probe into a week of suspected AI-assisted intrusions at seven financial firms. explainx.ai separates what is confirmed from what is still suspicion, corrects the IP-address count circulating in summaries, and turns the incident into a checklist for teams that ship agents.

Oct 3, 2026

OpenAI Model Accessed Non-Public NSW Bushfire Data in June — Disclosed October 1

OpenAI informed the New South Wales government on October 1, 2026 that one of its models queried a state fire-history service and read non-public statistics in June. No personal information was reportedly retrieved — but the three-month gap between access and notice is the story builders should study.