inferen-sh/skills▌
80 approved skills in this repository
social-media-carousel
Productivity
Multi-slide carousel design framework for Instagram, LinkedIn, Twitter/X with platform specs and engagement patterns. \n \n Provides a 7-slide structure (hook, context, 4 value slides, CTA) with templates for educational, storytelling, before/after, and listicle formats \n Includes platform-specific dimensions (Instagram/LinkedIn 1080x1350, Twitter/X up to 4 slides) and design rules for text hierarchy, readability, and visual consistency \n Covers swipe psychology principles: curiosity gaps, num
email-design
Frontend
High-converting email templates with layout patterns, subject line formulas, and mobile-first design rules. \n \n Covers five email types: welcome sequences, promotional campaigns, product updates, transactional receipts, and newsletters with specific templates and timing for each \n Enforces 600px max width, single-column layouts, 14-16px body text, and 44-50px CTA buttons optimized for 60%+ mobile opens \n Includes subject line formulas (numbered benefits, questions, personalization) capped at
video-prompting-guide
Frontend
Structured prompting techniques for generating high-quality videos across multiple AI platforms. \n \n Covers eight video generation models (Veo, Seedance, Wan, Grok, Kling, Runway, Pika, Sora) with model-specific optimization tips \n Provides a reusable prompt formula combining shot type, subject, action, setting, lighting, style, and technical parameters \n Includes reference tables for shot types, camera movements, lighting keywords, and visual aesthetics with practical examples \n Demonstrat
image-to-video
Video
Convert still images to animated videos with model selection, motion prompting, and camera control. \n \n Six models available (Wan 2.5 i2v, Seedance, Fabric, Grok) with guidance on when to use each based on content type and motion style \n Motion prompting framework covering camera movement (pan, dolly, orbit, crane), subject motion (natural elements, character, liquid), and atmospheric effects \n Best practices emphasizing subtle motion over dramatic action, with structured prompt templates an
technical-blog-writing
Productivity
Structured technical blog writing for developers with templates, code examples, and distribution guidance. \n \n Covers five post types: tutorials, deep dives, postmortems, benchmarks, and architecture posts, each with specific structure and word count targets \n Includes detailed rules for developer-friendly voice, code formatting, explanation depth, and what to avoid (filler language, dismissive tone, broken examples) \n Provides templates for ideal post structure, diagram generation via CLI,
app-store-screenshots
Productivity
Create platform-compliant app store screenshots and preview videos for iOS and Android. \n \n Generates screenshots at exact Apple App Store (iOS) and Google Play (Android) dimensions with device mockups, captions, and lifestyle contexts \n Covers critical gallery ordering strategy: first 3 screenshots drive 80% of impressions and must communicate core value, differentiation, and top features \n Includes preview video structure (15–30s for iOS, 30s–2min for Android) with hook, feature demonstrat
product-photography
Productivity
Professional product images with studio lighting, angles, and e-commerce conventions. \n \n Covers six shot types: hero shots, packshots (white background), lifestyle, scale reference, close-ups, and group compositions \n Includes camera angles (eye level, 3/4 view, overhead, low angle) and lighting setups (soft box, rim lighting, natural window, flat/even) \n Provides shadow and background guidance, composition rules (rule of thirds, negative space, triangle layouts), and category-specific conv
case-study-writing
Productivity
Structured B2B case study creation with STAR framework, metrics visualization, and research integration. \n \n Follows the Situation-Task-Action-Result framework with templates for headline, snapshot box, challenge, solution, and results sections \n Emphasizes quantified metrics and before/after comparisons across time, money, efficiency, growth, and satisfaction categories \n Includes guidance on customer quotes, data visualization via Python charts, and industry research using search tools \n
infsh-cli
Productivity
Access 150+ cloud-based AI apps via CLI without GPU setup or infrastructure. \n \n Covers image generation (FLUX, Gemini, Grok), video creation (Veo, Seedance, OmniHuman), LLMs (Claude, Gemini via OpenRouter), web search (Tavily, Exa), 3D modeling, and Twitter/X automation \n Automatically uploads local files when provided as paths; supports batch operations and async execution with task status polling \n Simple command structure: infsh app list , infsh app run <app> --input data.json , wit
nano-banana-2
Productivity
Text-to-image and image editing with Google Gemini 3.1 Flash, supporting up to 14 input images and real-time search grounding. \n \n Generates single or multiple images with customizable aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4) and resolutions up to 4K \n Edits existing images by accepting up to 14 input images (JPEG, PNG, WebP) alongside a text prompt \n Enables Google Search grounding to incorporate real-time information like weather and news into generated imagery \n Accessible via inference
nano-banana
Productivity
Generate images with Google Gemini native image models via inference.sh CLI. \n \n Two models available: Gemini 3 Pro Image (highest quality, slower) and Gemini 2.5 Flash Image (fast, excellent quality) \n Supports text-to-image generation, image editing with up to 14 input images, custom aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4), and resolutions up to 4K \n Includes optional Google Search grounding for real-time information integration and batch generation of multiple images per prompt \n Requi
landing-page-design
Frontend
Landing page conversion optimization with layout rules, hero design, and CTA psychology. \n \n Covers above-the-fold formula, headline patterns, hero image strategy, and section sequencing proven to convert \n Includes CTA button design rules, form field optimization, mobile-first guidelines, and social proof placement strategies \n Provides specific examples of high-converting vs. failing headlines, CTAs, and hero images with anti-patterns \n Integrates with inference.sh CLI to generate hero im
ai-image-generation
AI/ML
Generate images with 50+ AI models including FLUX, Gemini, Grok, and Seedream via inference.sh CLI. \n \n Supports text-to-image, image-to-image, inpainting, LoRA customization, image editing, upscaling, and text rendering across multiple model families \n Models range from ultra-fast budget options (FLUX Klein at $0.0001/image) to high-fidelity 4K outputs (Seedream 4.5, ImagineArt 1.5 Pro) \n Includes Google Gemini, xAI Grok, ByteDance Seedream, and Pruna P-Image variants with configurable aspe
storyboard-creation
Productivity
Visual storyboarding with shot vocabulary, camera angles, continuity rules, and AI image generation via inference.sh. \n \n Covers eight shot types (ECU to EWS), seven camera angles (eye level to Dutch), and nine camera movements (pan, dolly, crane, handheld, etc.) with use cases and emotional effects \n Enforces continuity best practices: 180-degree rule, match on action, eyeline match, and screen direction to maintain spatial coherence across cuts \n Generates storyboard panels using FLUX imag
image-upscaling
Productivity
Upscale and enhance images using Topaz, Real-ESRGAN, and other inference.sh models. \n \n Supports multiple upscaling models including Topaz Image Upscaler, Real-ESRGAN, and FLUX Dev Upscaler via the inference.sh CLI \n Handles common image enhancement tasks: restoring old photos, preparing AI-generated art for print, creating high-resolution versions for displays, and enlarging thumbnails \n Integrates with image generation workflows, allowing you to generate and upscale images in sequence \n R
book-cover-design
Frontend
Genre-specific book cover design with AI image generation, typography guidance, and platform sizing. \n \n Covers fiction (thriller, romance, sci-fi, fantasy, literary, horror, historical) and non-fiction (business, memoir, science, cookbook, travel) with genre-specific color palettes, imagery conventions, and typography pairings \n Includes print trim sizes (mass market to large format), digital platform specs (Kindle, Apple Books, general ebook), and spine width calculations \n Provides the th
ai-voice-cloning
AI/ML
Natural AI voice generation across seven models with 22+ voices, multiple languages, and emotional range. \n \n Supports ElevenLabs (premium quality, 32 languages), Kokoro TTS, DIA, Chatterbox, Higgs, and VibeVoice, each optimized for different styles from professional narration to casual conversation \n Includes 16+ named voices with gender and style profiles (e.g., warm, authoritative, youthful) plus speed control (0.8–1.2x) and punctuation-based pacing \n Handles multi-voice conversations, lo
related-skill
Productivity
Discover and install complementary skills from the inference.sh registry to expand your AI agent's capabilities. \n \n Browse 150+ available skills across media generation, image processing, audio, search, and social automation categories \n Search, list, and install skills using simple CLI commands ( npx skills search , npx skills add ) \n Pre-configured skill combinations for common workflows: research agents, content creators, and media processors \n Manage installed skills with built-in comm
ai-social-media-content
AI/ML
Generate images, videos, captions, and thumbnails for TikTok, Instagram, YouTube, and Twitter/X. \n \n Supports all major platform formats: vertical 9:16 for TikTok/Reels/Shorts, 16:9 for YouTube thumbnails and Twitter, and 1:1 for Instagram Feed \n Integrates video generation (Veo, Seedance), image creation (FLUX), text-to-speech (Kokoro), and AI avatars for talking head content \n Includes workflow templates for trending videos, tutorials, product showcases, lifestyle content, and behind-the-s
python-executor
Backend
Execute arbitrary Python code in a sandboxed environment with 100+ pre-installed libraries. \n \n Supports data processing (NumPy, Pandas, SciPy), web scraping (requests, BeautifulSoup, Selenium, Playwright), and image/video manipulation (Pillow, OpenCV, MoviePy) \n Includes 3D model processing (trimesh, open3d), PDF generation (reportlab, pypdf2), and SVG creation (svgwrite, cairosvg) \n Configurable timeout (1–300 seconds), output capture, and working directory; optional high-memory variant (1
elevenlabs-tts
Productivity
Premium text-to-speech with 22+ voices, three quality tiers, and multilingual support across 32 languages. \n \n Three models available: eleven_multilingual_v2 for highest quality, eleven_turbo_v2_5 for balanced speed, and eleven_flash_v2_5 for ultra-low latency (~75ms) \n 22+ named voices across male and female personas with distinct styles (e.g., aria for conversational American, george for authoritative British) \n Fine-tune voice output via stability, similarity_boost, and style parameters;
javascript-sdk
Backend
JavaScript/TypeScript SDK for running AI apps, building agents, and integrating 150+ models. \n \n Supports running AI apps with basic execution, fire-and-forget, and streaming progress modes; includes automatic file upload and stateful sessions to keep workers warm across calls \n Agent SDK enables both template-based agents from your workspace and ad-hoc agents with custom tools, system prompts, and temperature control \n Tool builder API provides four tool types: client tools (run in your cod
web-search
Productivity
Web search and content extraction via Tavily and Exa APIs through inference.sh CLI. \n \n Five search and extraction apps: Tavily Search Assistant (AI-powered answers with sources), Tavily Extract (multi-URL content extraction), Exa Search (smart web search), Exa Answer (direct factual responses), and Exa Extract (web page analysis) \n Designed for research workflows, RAG pipelines, fact-checking, and content aggregation with LLM integration examples \n Requires inference.sh CLI ( infsh ) instal
elevenlabs-music
Productivity
Generate original music from text prompts with customizable duration, genre, and mood control. \n \n Supports durations from 5 seconds to 10 minutes, enabling use cases from notification sounds to full cinematic scores \n Text prompts accept genre, mood, instrument, tempo, and style descriptors for fine-grained control over output \n Royalty-free music suitable for commercial use in videos, podcasts, games, ads, and presentations \n Integrates with ElevenLabs TTS and sound effects skills for com
p-image
Productivity
Fast, optimized image generation with Pruna's P-Image models via inference.sh CLI. \n \n Four model variants: P-Image for text-to-image, P-Image-LoRA with 11 preset styles, P-Image-Edit for image editing, and P-Image-Edit-LoRA for stylized edits \n Supports multiple aspect ratios (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, custom) and multi-image compositing for collages and combinations \n Requires inference.sh CLI ( infsh ) and login; run models via infsh app run pruna/[model-name] with JSON input p
elevenlabs-dialogue
Productivity
Generate multi-speaker dialogue audio with 22+ premium voices and voice direction controls. \n \n Supports 8 female and 14 male voices with preset pairings optimized for interviews, tutorials, podcasts, and debates \n Voice direction keywords (excitedly, sadly, whispering, angrily, sarcastically, etc.) control delivery tone and emotion for each segment \n Input format uses JSON segments with text and voice assignment; output is a single audio file combining all speakers \n Designed for podcasts,
building-inferencesh-apps
Frontend
Build and deploy applications on the inference.sh platform. Apps can be written in Python or Node.js.
elevenlabs-voice-isolator
Productivity
Remove background noise and isolate vocals from audio files via inference.sh CLI. \n \n Supports five audio formats (WAV, MP3, FLAC, OGG, AAC) up to 500MB and 1 hour duration \n Removes ambient noise, background music, reverb, wind, traffic, electrical hum, and other non-voice sounds \n Integrates into multi-step workflows for transcription, voice transformation, and video production \n Requires inference.sh CLI ( infsh ) installation and login to run \n
pitch-deck-visuals
Productivity
Investor-ready pitch deck framework with 12-slide structure, visual design rules, and chart guidelines. \n \n Provides a battle-tested 12-slide framework covering problem, solution, market size, traction, team, and financials, with recommended timing for each slide \n Includes typography, color, and layout rules designed for investor clarity: the 1-6-6 rule (one idea, six words, six bullets per slide), dark or clean backgrounds, and consistent margins \n Specifies which chart types work for each
elevenlabs-dubbing
Productivity
Translate and dub audio/video into 29 languages while preserving original speaker voice. \n \n Supports 29 languages including Spanish, French, German, Chinese, Japanese, Arabic, and more; accepts MP3, MP4, WAV, and MOV files \n Automatically detects speakers and source language, or accepts explicit language specification for precision \n Handles multi-speaker videos while maintaining original voice characteristics, timing, and pacing \n Integrates with inference.sh CLI for single-command dubbin
logo-design-guide
Frontend
AI-powered logo design with prompting strategies, scalability rules, and iteration workflows. \n \n Covers six logo types (wordmark, lettermark, pictorial, abstract, mascot, combination) with use cases and design considerations for each \n Provides tested prompt structures and keywords that work with AI image generators, plus explicit anti-patterns to avoid (text rendering, photorealism, gradients) \n Includes scalability checklist ensuring logos work from 16px favicons to billboards, plus color
elevenlabs-voice-changer
Productivity
Transform any voice into a different voice while preserving speech content and emotion. \n \n Supports two models: multilingual STS v2 (70+ languages) and English-optimized STS v2 \n Access to 22+ premium voices across multiple styles and accents (British, American, Australian, and more) \n Configurable output formats and seamless integration with media workflows for voice-over replacement and character creation \n Requires inference.sh CLI ( infsh ) for command-line execution \n
dialogue-audio
Productivity
Realistic multi-speaker dialogue audio generation with Dia TTS via inference.sh CLI. \n \n Supports two-speaker conversations with automatic voice assignment using [S1] and [S2] speaker tags \n Emotion and pacing controlled through punctuation ( . , ! , ? , ... , — ) and parenthetical sound cues like (laughs) , (sighs) , and (whispers) \n Includes structured patterns for interviews, tutorials, debates, and conversational content with practical script-writing guidelines \n Post-production support
python-sdk
Backend
Python SDK for building AI applications, agents, and integrations with 150+ inference.sh models. \n \n Supports sync and async execution with streaming progress updates, fire-and-forget task submission, and stateful sessions for warm worker continuity \n Agent SDK includes template agents, ad-hoc agent creation with custom tools, and built-in capabilities like web search, code execution, and image generation \n Tool builder API enables four tool types: client tools (run in your code), app tools
widgets-ui
Frontend
Render rich interactive UIs from JSON-structured widget definitions in React/Next.js. \n \n Supports 15+ widget types across layouts (row, col, box), typography (title, text, caption), interactive elements (button, input, select, checkbox), and display components (badge, icon, image, divider) \n Includes five preset gradients (ocean, sunset, purple, cool, midnight) for styled containers and flexible spacing/alignment options \n Handles form submission and click actions through a unified action h
newsletter-curation
Productivity
Structured newsletter creation with content sourcing, curation templates, and subscriber growth strategies. \n \n Provides five newsletter formats (link roundup, deep dive, essay, Q&A, data/trends) with ready-to-use markdown templates and issue structure guidance \n Includes content sourcing via Tavily and Exa search integration, plus a curation quality filter to evaluate pieces before inclusion \n Covers writing commentary formulas, subject line strategies, sending cadence optimization (we
agent-tools
Productivity
Access 150+ cloud-hosted AI apps via CLI without GPU requirements. \n \n Covers image generation (FLUX, Gemini, Grok), video creation (Veo, Seedance, OmniHuman), LLMs (Claude, Gemini via OpenRouter), web search (Tavily, Exa), 3D modeling, and Twitter/X automation \n Automatically uploads local files when provided as paths, enabling workflows with images, audio, and media without manual URL conversion \n Run apps synchronously or asynchronously with --no-wait flag; check task status by ID for lon
ai-product-photography
AI/ML
Professional product photography generation across multiple AI models with studio, lifestyle, and e-commerce templates. \n \n Supports five image generation models (FLUX Dev/Schnell, Imagen 3, Grok, Seedream) optimized for different quality and speed tradeoffs \n Includes pre-built templates for common product photography styles: studio white background, flat lay, hero shots, lifestyle context, and in-use action shots \n Covers six product categories with category-specific prompts: electronics,
video-ad-specs
Video
Platform-specific video ad dimensions, durations, and creative frameworks for TikTok, Instagram, YouTube, Facebook, and LinkedIn. \n \n Covers exact specs for five platforms: TikTok and Instagram (9:16 vertical), YouTube (6s bumper to skippable formats), Facebook (1:1 or 4:5 square), and LinkedIn (1:1 or 16:9) \n Includes AIDA framework (Attention, Interest, Desire, Action) with timing breakdowns and hook techniques for the critical first 3 seconds \n Provides caption requirements and safe-zone
ai-rag-pipeline
AI/ML
Build RAG pipelines combining web search and LLMs for grounded, sourced AI responses. \n \n Integrates multiple search tools (Tavily, Exa) and LLM providers (Claude, GPT-4, Gemini via OpenRouter) via the inference.sh CLI \n Supports three core patterns: simple search-plus-answer, multi-source research aggregation, and URL content extraction with analysis \n Includes ready-to-use examples for fact-checking, research reports, and iterative deep-dive queries with built-in source attribution \n Best
customer-persona
Productivity
Research-backed customer personas with demographics, psychographics, jobs-to-be-done, and avatar generation. \n \n Provides a structured template covering demographics, psychographics, goals, pain points, buying triggers, and anti-personas for market segmentation \n Integrates with inference.sh CLI to research market data via Tavily and Exa search, then generate realistic persona avatars using Flux image generation \n Includes step-by-step guidance for validating personas through customer interv
character-design-sheet
Frontend
Maintain consistent character appearance across AI-generated images using reference sheets and detailed descriptions. \n \n Create turnaround sheets (front, 3/4, side, back views), expression sheets (6+ emotions), outfit variations, and color palettes to document character design \n Use a 50+ word detailed description anchor reused exactly in every prompt as the most practical consistency technique for small projects \n Train character-specific LoRA models for projects requiring many images, wit
flux-image
Productivity
Text-to-image and image-to-image generation using FLUX models via inference.sh CLI. \n \n Supports multiple FLUX variants: Dev LoRA (highest quality), Klein LoRA (fastest), and Pruna-optimized versions ranging from 4B to full-size models \n Enables LoRA fine-tuning for custom style adaptation and image-to-image transformations alongside standard text-to-image generation \n Requires inference.sh CLI ( infsh ) installation and authentication; models are invoked as named apps with JSON input payloa
explainer-video-guide
Frontend
Complete guide for scripting, producing, and assembling explainer videos via inference.sh CLI. \n \n Provides three script formulas (Problem-Agitate-Solve, Before-After-Bridge, Feature Spotlight) with timing breakdowns, word counts, and pacing rules for 30–120 second videos \n Covers scene generation via video AI models, voiceover production with TTS, music ducking, and caption integration through a multi-step assembly pipeline \n Includes practical tables for pacing (120–170 wpm by content type
content-repurposing
Marketing
Atomize long-form content into 10+ derivative formats across platforms. \n \n Covers six core conversions: blog-to-thread, blog-to-carousel, blog-to-newsletter, podcast-to-blog, video-to-clips, and quote-card generation \n Includes platform-specific adaptation rules for Twitter, LinkedIn, TikTok, email, and short-form video, accounting for tone, attention span, and format constraints \n Provides conversion recipes with concrete specs (tweet length, slide counts, video duration) and example CLI c
ai-music-generation
AI/ML
Generate music and songs via three AI models through the inference.sh CLI. \n \n Three models available: ElevenLabs Music (up to 10 min with commercial license), Diffrythm (fast generation), and Tencent Song Generation (full songs with vocals) \n Supports text-to-music, instrumental tracks, lyrics-to-song, and soundtrack creation via simple CLI commands \n Prompt-based generation using genre, mood, instrument, and structure keywords for precise control \n Ideal for social media content, podcasts
talking-head-production
Productivity
AI avatar talking head videos with lipsync, TTS, and multi-character support via inference.sh CLI. \n \n Supports OmniHuman 1.5 for multi-character conversations and gestures, OmniHuman 1.0 for single characters, and PixVerse for quick lipsync on existing video \n Requires high-quality source portraits (min 512x512, ideally 1024x1024+) with frontal gaze, neutral expression, and head-and-shoulders framing for accurate animation \n Generates dialogue audio via Dia TTS with speaker tags and emotion
twitter-thread-creation
Productivity
Craft high-engagement Twitter/X threads with optimized hooks, structure, and formatting. \n \n Provides hook templates (bold claim, contrarian take, story opener, how-to promise) and a proven thread anatomy: hook tweet, context, numbered content points (one idea per tweet), summary, and call-to-action \n Includes character limits for free (280) and Premium (25,000) accounts, image specifications (1200×675 recommended), and formatting rules using symbols (→, •, ✅, ❌) for clarity and pacing \n Cov
press-release-writing
Productivity
Professional press releases in AP style with inverted pyramid structure and fact-checking. \n \n Covers all major release types: product launches, funding announcements, partnerships, milestones, and executive hires \n Includes detailed AP style formatting rules, headline guidelines, dateline conventions, and quote attribution standards \n Provides research workflows using inference.sh CLI to verify claims, check market data, and gather industry context before writing \n Features section-by-sect
product-changelog
Productivity
Changelogs and release notes that users actually read and act on. \n \n Covers entry structure, user-facing language, and categorization (New, Improved, Fixed, Removed, Security) with clear rules for when to use each \n Includes semantic and date-based versioning strategies, visual changelog guidance, and templates for breaking changes with migration timelines \n Provides distribution channel recommendations (changelog page, in-app, email, blog, social) and writing best practices to avoid common
ai-content-pipeline
AI/ML
Multi-step AI content creation pipelines combining image, video, audio, and text generation. \n \n Chains tools like FLUX (image generation), Wan 2.5 (animation), Kokoro TTS (voice), and OmniHuman (talking heads) into complete workflows via the inference.sh CLI \n Includes four ready-to-use pipeline templates: YouTube shorts (script → voiceover → background → animation → merge), talking head videos, product demos, and blog-to-video conversion \n Supports common patterns such as image-to-video-to
google-veo
Backend
Text-to-video generation using Google Veo models via inference.sh CLI. \n \n Five Veo model variants available, ranging from Veo 2 to Veo 3.1, with trade-offs between speed and quality \n Supports cinematic, product, nature, action, and urban scene generation with detailed prompt control over camera movements, lighting, and timing \n Requires inference.sh CLI ( infsh ) for authentication and app execution \n Includes sample workflow for generating input templates and iterating on prompts \n
background-removal
Productivity
Remove backgrounds from images via inference.sh CLI, returning transparent PNGs. \n \n Uses BiRefNet model for high-accuracy background removal on product photos, portraits, and general images \n Returns PNG files with transparent backgrounds ready for e-commerce, design, and social media use \n Integrates with Reve for advanced editing workflows, including background replacement and direct prompt-based modifications \n Chainable with image generation and upscaling skills for complete photo edit
text-to-speech
Productivity
Multiple text-to-speech models via inference.sh CLI for voiceovers, podcasts, and accessibility. \n \n Six models available: ElevenLabs (premium, 22+ voices, 32 languages), DIA TTS (conversational), Kokoro TTS (fast), Chatterbox, Higgs Audio (emotional control), and VibeVoice (long-form podcasts) \n Core capabilities include basic speech synthesis, expressive speech with emotion control, and conversational dialogue generation \n Easily combine with video tools like OmniHuman to create talking he
youtube-thumbnail-design
Frontend
AI-generated YouTube thumbnails optimized for mobile preview and click-through rates. \n \n Requires 1280×720px minimum (1920×1080px recommended) with high contrast color pairs and max 3 colors per thumbnail \n Includes the 120px mobile test: thumbnail must clearly show mood, subject, and readable text when viewed at that width \n Safe zone guidelines prevent critical elements from being obscured by video duration timestamps (bottom-right) and chapter markers (bottom-left) \n Face expressions si
ai-avatar-video
AI/ML
Generate talking head and avatar videos from images and audio using OmniHuman, Fabric, and PixVerse models. \n \n Four model options: OmniHuman 1.5 (multi-character), OmniHuman 1.0 (single character), Fabric 1.0 (image lipsync), and PixVerse Lipsync (highly realistic) \n Audio-driven workflow: pair portrait images with speech files to generate realistic avatar videos with synchronized lip movement \n Composable with text-to-speech and video transcription for end-to-end pipelines: generate speech
og-image-design
Frontend
Design and generate Open Graph images optimized for social platform sharing with platform-specific dimensions and text placement. \n \n Covers seven major platforms (Facebook, Twitter/X, LinkedIn, Discord, Slack, iMessage) with standardized 1200 x 630 px dimensions and safe zone padding rules \n Includes pre-built HTML templates for blog posts, product launches, and tutorials, executable via the infsh html-to-image command \n Provides design guidelines for typography (48–64px titles, 20–28px sub
competitor-teardown
Productivity
Structured competitive analysis with feature matrices, pricing breakdowns, SWOT analysis, and positioning maps. \n \n Provides a 7-layer analysis framework covering product, pricing, positioning, traction, reviews, content, and team across competitors \n Includes templates and commands for feature matrices, pricing comparisons, SWOT analysis, and 2x2 positioning maps with visual generation \n Integrates web search (Tavily, Exa), browser automation for UX screenshots, and Python for positioning m
prompt-engineering
Productivity
Techniques and patterns for crafting effective prompts across LLMs, image generators, and video models. \n \n Covers LLM prompting fundamentals: role assignment, task clarity, chain-of-thought reasoning, few-shot examples, output format specification, and constraint setting \n Image generation structure includes subject description, style keywords, composition control, quality modifiers, and negative prompt usage \n Video prompting guidance covers shot types, camera movement, action description,
tools-ui
Frontend
React/Next.js components for displaying tool calls across their full lifecycle: pending, running, approval, success, and error states. \n \n Includes three core components: ToolCall for displaying pending/running tool invocations, ToolResult for showing completed outputs, and ToolApproval for human-in-the-loop approval flows with approve/deny callbacks \n Automatic icon assignment based on tool name patterns (search, read, write, delete, send, etc.) with fallback to wrench icon \n Five distinct
data-visualization
Productivity
Clear, effective data visualizations with chart selection rules, design principles, and storytelling techniques. \n \n Covers 10+ chart types with decision rules for when to use each (line for time series, bar for comparison, scatter for correlation, heatmap for patterns) \n Design guidelines for axes, color theory, typography, and annotations including colorblind-safe palettes and a strong stance against pie charts \n Includes ready-to-run Python/matplotlib recipes for line charts, bar charts,
product-hunt-launch
Productivity
Optimize Product Hunt launches with research, gallery strategy, and timing guidance. \n \n Covers listing specifications (tagline limits, gallery dimensions, topic selection) and gallery image positioning strategy with five-image framework from hero shot to social proof \n Provides tagline formulas and examples to communicate product value in 60 characters without buzzwords \n Includes launch day playbook with specific timing (12:01 AM PT), maker comment structure, and engagement timeline from p
llm-models
AI/ML
Access 100+ language models including Claude, Gemini, Kimi, and GLM via OpenRouter. \n \n Supports Claude Opus 4.5, Sonnet 4.5, and Haiku 4.5 for varying performance and cost tradeoffs, plus Gemini 3 Pro, Kimi K2 Thinking, GLM-4.6, and Intellect 3 \n Auto-select mode picks the most cost-effective model automatically for your use case \n Designed for code generation, reasoning, content creation, data analysis, and building AI agent workflows \n Requires inference.sh CLI ( infsh ) and API authenti
agent-browser
Productivity
Playwright-based browser automation with element refs for AI agents, supporting navigation, interaction, screenshots, and video recording. \n \n Provides 6 core functions: open (navigate with config), snapshot (refresh element refs), interact (click/fill/drag/upload/scroll), screenshot, execute (JavaScript), and close \n Element interaction uses simple @e ref system that invalidates after navigation, requiring re-snapshot calls to maintain accurate selectors \n Supports video recording with opti
ai-automation-workflows
AI/ML
Automate AI workflows combining multiple models and services via bash scripting and the inference.sh CLI. \n \n Supports five core patterns: batch processing, sequential pipelines, parallel execution, conditional branching, and retry logic with fallbacks \n Integrates with inference.sh CLI for model invocation, bash for orchestration, and Python SDK for programmatic workflows \n Includes cron job setup for scheduled automation, logging wrappers for monitoring, and webhook-based error alerting \n
chat-ui
Frontend
React/Next.js chat UI components for building custom messaging interfaces and AI assistant conversations. \n \n Includes four core components: ChatContainer for layout, ChatMessage for user/assistant/system messages, ChatInput for message submission, and TypingIndicator for loading states \n Supports role-based message variants (user, assistant, system) with automatic alignment and styling \n Built on Tailwind CSS and shadcn/ui design tokens with customizable className props \n Installable via s
agent-ui
Frontend
Drop-in React/Next.js agent component with runtime, tools, streaming, and human-in-the-loop approvals built in. \n \n Single component handles agent execution, tool lifecycle management, real-time token streaming, and approval flows without backend logic \n Supports client-side tools that run in the browser, file and image uploads, and declarative JSON widgets for agent-generated UI \n Includes API proxy route setup for Next.js and configurable agent parameters like model selection, system promp
remotion-render
Video
Convert React/Remotion component code directly to MP4 videos via CLI. \n \n Accepts TSX code with full Remotion API support: useCurrentFrame , useVideoConfig , spring , interpolate , AbsoluteFill , Sequence , and media components \n Configurable output: resolution, FPS, duration, codec, and component props passed as JSON \n Supports streaming progress updates and Python SDK integration for programmatic video generation \n Ideal for animated graphics, motion design, data-driven videos, and React
qwen-image-2-pro
Productivity
Professional image generation with advanced text rendering, photorealistic details, and semantic accuracy via Alibaba Qwen-Image-2.0-Pro. \n \n Excels at text-heavy designs including posters, banners, and multi-line text with fine-grained control over font, color, and positioning \n Generates 1–6 images per request with configurable dimensions (512–2048 pixels), negative prompts, and reproducible seeds \n Supports image editing and style transfer via reference images for outfit swaps and creativ
p-video
Video
Generate videos from text or images using Pruna's optimized models via inference.sh CLI. \n \n Three models available: P-Video (text-to-video, image-to-video, audio support), WAN-T2V (fast text-to-video), and WAN-I2V (image animation) \n Supports 480p, 720p, and 1080p resolutions with draft mode for faster, cheaper generation \n Audio synchronization available for P-Video to create videos with synced audio tracks \n Requires inference.sh CLI ( infsh ) and authentication via infsh login \n
qwen-image-2
Productivity
Text-to-image and multi-image editing with Alibaba Qwen-Image-2.0 models via inference.sh CLI. \n \n Two models available: Qwen-Image-2.0 for fast general use, and Qwen-Image-2.0-Pro for professional text rendering and fine-grained control \n Supports text-to-image generation, multi-reference image editing (up to 3 input images), custom resolutions (512–2048 pixels), and negative prompts \n Key parameters include prompt extension toggle, seed-based reproducibility, watermark control, and batch g
twitter-automation
Productivity
Post, like, retweet, and manage Twitter/X accounts via CLI commands. \n \n Nine app commands covering tweet posting with media, liking, retweeting, DMs, user following, and profile retrieval \n Integrates with inference.sh CLI ( infsh ) for direct command-line automation of X/Twitter actions \n Supports media attachment workflows by chaining with image and video generation apps \n Requires inference.sh CLI installation and X API authentication via infsh login \n
linkedin-content
Marketing
Write high-engagement LinkedIn posts with proven hook formulas, formatting rules, and algorithm optimization. \n \n Covers seven hook formulas (contrarian opinions, personal stories, surprising stats, lists, bold statements, before/after, pattern interrupts) and identifies common failures like corporate jargon and weak openings \n Post anatomy guidance includes the critical 210-character hook visible before \"see more,\" body formatting with line breaks for mobile readability, and CTA strategies
speech-to-text
Productivity
Transcribe audio to text using ElevenLabs Scribe or Whisper models via inference.sh CLI. \n \n Three model options: ElevenLabs Scribe v2 (98%+ accuracy with diarization), Fast Whisper V3, and Whisper V3 Large for varying speed/accuracy tradeoffs \n Supports 99+ languages, optional timestamps, speaker diarization, and translation to English \n Common workflows include meeting transcription, podcast transcripts, video subtitles, and voice note conversion \n Requires inference.sh CLI ( infsh ) inst
ai-marketing-videos
AI/ML
Generate professional marketing videos for ads, promos, product launches, and brand content. \n \n Supports multiple video generation models (Veo, Seedance, Wan, FLUX) and Kokoro for AI voiceovers, enabling end-to-end video creation from prompt to final asset \n Includes templates and workflows for common ad types: 6–90 second bumper ads, product demos, testimonials, explainers, and before/after transformations \n Provides platform-specific guidance for Facebook, Instagram, YouTube, TikTok, and
elevenlabs-stt
Productivity
98%+ accurate transcription with speaker diarization, audio event tagging, and word-level forced alignment. \n \n Supports Scribe v1 and v2 models with auto-detection across 90+ languages \n Capabilities include speaker identification, audio event tagging (laughter, applause, music), and precise word-level timestamps via forced alignment \n Forced alignment enables subtitle generation, lip-sync timing, and karaoke applications by aligning known text to audio \n Requires inference.sh CLI ( infsh
elevenlabs-sound-effects
Productivity
Generate royalty-free sound effects from text descriptions using ElevenLabs AI. \n \n Supports custom duration (0.5–22 seconds) and prompt influence control (0–1) to balance creative interpretation against literal accuracy \n Three parameter types: text description, optional duration, and optional prompt influence for fine-tuning output \n Covers cinematic effects, nature ambience, game audio, and everyday sounds with practical examples for each category \n Integrates with inference.sh CLI and p
ai-podcast-creation
AI/ML
Multi-voice podcast and audiobook production with TTS, AI music, and audio merging. \n \n Supports three TTS engines (Kokoro, DIA, Chatterbox) with 6+ voice options across American, British, and conversational styles \n Includes AI music generation for intros, outros, and background tracks; media merger handles crossfades and layering \n Workflows cover single narration, multi-voice dialogue, full episode pipelines, and NotebookLM-style document discussions \n Integrates with Claude for script g
seo-content-brief
Marketing
Data-driven SEO content briefs with keyword research, SERP analysis, and content structure planning. \n \n Provides a structured template covering target keywords, search intent analysis, competitor gaps, heading hierarchy, and word count targets matched to top-ranking results \n Includes SERP analysis workflows using inference.sh CLI to extract competitor data, identify content format patterns, and find ranking opportunities via \"People Also Ask\" \n Covers keyword clustering, on-page SEO chec
ai-video-generation
AI/ML
Generate videos with 40+ AI models including Veo, Seedance, Wan, and Grok via inference.sh CLI. \n \n Supports text-to-video, image-to-video, avatar animation, lipsync, video upscaling, and foley sound generation across multiple model families \n Access 10+ text-to-video models (Veo 3.1, Seedance 1.5 Pro, Wan, Grok Video) and 5+ image-to-video variants optimized for speed, quality, or cost \n Includes avatar and lipsync tools (OmniHuman, Fabric, PixVerse) for talking-head and character animation