Back to directory
rafska avatar
rafska / awesome-local-llm

awesome-local-llm

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

2,513

Stars

316

Forks

32

Watchers

MIT

License

Awesome local LLM badge

A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally

Table of Contents

Inference platforms

  • LM Studio - discover, download and run local LLMs
  • unsloth unsloth - unified web UI for training and running open models like Qwen, DeepSeek, and Gemma locally
  • LocalAI LocalAI - the free, open-source alternative to OpenAI, Claude and others
  • jan jan - an open source alternative to ChatGPT that runs 100% offline on your computer
  • ChatBox ChatBox - user-friendly desktop client app for AI models/LLMs
  • lemonade lemonade - a local LLM server with GPU and NPU Acceleration

Back to Table of Contents

Inference engines

  • ollama ollama - get up and running with LLMs
  • llama llama.cpp - LLM inference in C/C++
  • vllm vllm - a high-throughput and memory-efficient inference and serving engine for LLMs
  • exo exo - run your own AI cluster at home with everyday devices
  • BitNet BitNet - official inference framework for 1-bit LLMs
  • sglang sglang - a fast serving framework for large language models and vision language models
  • NVIDIA 25 TensorRT LLM TensorRT-LLM - provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs
  • nano vllm Nano-vLLM - a lightweight vLLM implementation built from scratch
  • omlx omlx - LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
  • koboldcpp koboldcpp - run GGUF models easily with a KoboldAI UI
  • mistral mistral.rs - fast, flexible LLM inference
  • NVIDIA 25 dynamo dynamo - a datacenter scale distributed inference serving framework
  • flashinfer flashinfer - kernel library for LLM serving
  • mlx lm mlx-lm - generate text and fine-tune large language models on Apple silicon with MLX
  • gpustack gpustack - simple, scalable AI model deployment on GPU clusters
  • Google 234285F4 LiteRT LM LiteRT-LM - Google's production-ready, high-performance, open-source inference framework for deploying Large Language Models on edge devices
  • mlx vlm mlx-vlm - a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX
  • executorch executorch - on-device AI across mobile, embedded and edge for PyTorch
  • mini sglang mini-sglang - a lightweight yet high-performance inference framework for Large Language Models
  • distributed llama distributed-llama - connect home devices into a powerful cluster to accelerate LLM inference
  • Google 234285F4 litert LiteRT - Google's on-device framework for high-performance ML & GenAI deployment on edge platforms, via efficient conversion, runtime, and optimization
  • ik llama ik_llama.cpp - llama.cpp fork with additional SOTA quants and improved performance
  • sonar sonar - large-scale LLM inference engine based on vLLM
  • FastFlowLM FastFlowLM - run LLMs on AMD Ryzen™ AI NPUs
  • tokenspeed tokenspeed - a speed-of-light LLM inference engine
  • krasis krasis - a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
  • vllm gfx906 vllm-gfx906 - vLLM for AMD gfx906 GPUs, e.g. Radeon VII / MI50 / MI60
  • llm scaler llm-scaler - run LLMs on Intel Arc™ Pro B60 GPUs

Back to Table of Contents

User Interfaces

  • open webui Open WebUI - User-friendly AI Interface (Supports Ollama, OpenAI API, ...)
  • lobe chat Lobe Chat - an open-source, modern design AI chat framework
  • text generation webui Text generation web UI - LLM UI with advanced features, easy setup, and multiple backend support
  • SillyTavern SillyTavern - LLM Frontend for Power Users
  • page assist Page Assist - Use your locally running AI models to assist you in your web browsing

Back to Table of Contents

Large Language Models

Explorers, Benchmarks, Leaderboards

  • Arena - benchmark & compare the best AI models
  • AI Models & API Providers Analysis - understand the AI landscape to choose the best model and provider for your use case
  • SWE-rebench - a continuously evolving and decontaminated benchmark for software engineering LLMs
  • bullshit benchmark BullshitBench - measure whether AI models challenge nonsensical prompts instead of confidently answering them
  • LLM Explorer - explore list of the open-source LLM models
  • Dubesor LLM Benchmark table - small-scale manual performance comparison benchmark
  • oobabooga benchmark - a list sorted by size (on disk) for each score
  • CyberGym - evaluating AI agents' real-world cybersecurity capabilities at scale
  • vakra vakra - a benchmark for evaluating multi-hop, multi-source tool-calling in AI agents

Back to Table of Contents

Model providers

  • Qwen - powered by Alibaba Cloud
  • Mistral 20AI 23FA520F Mistral AI - a pioneering French artificial intelligence startup
  • Tencent - a profile of a Chinese multinational technology conglomerate and holding company
  • Unsloth AI - focusing on making AI more accessible to everyone (GGUFs etc.)
  • bartowski - providing GGUF versions of popular LLMs
  • Beijing Academy of Artificial Intelligence - a private non-profit organization engaged in AI research and development
  • Open Thoughts - a team of researchers and engineers curating the best open reasoning datasets

Back to Table of Contents

Specific models

General purpose

  • DeepSeek-V4 - a collection of the DeepSeek V4 LLMs
  • Qwen3.6 - a collection of the latest generation Qwen LLMs
  • NVIDIA 25 NVIDIA Nemotron v3 - a family of open models from NVIDIA with open weights, training data and recipes, delivering leading efficiency and accuracy for building specialized AI agents
  • Google 234285F4 Gemma 4 - a family of open models built by Google DeepMind, that are multimodal, handling text and image input (with audio supported on small models) and generating text output
  • Mistral 20AI 23FA520F Mistral Medium 3.5 - The first flaship models from Mistral AI handling instruction-following, reasoning, and coding in a single set of opened-weights
  • OpenAI 23412991 gpt-oss - a collection of open-weight models from OpenAI, designed for powerful reasoning, agentic tasks, and versatile developer use cases
  • NVIDIA 25 gpt-oss-puzzle-88B - a deployment-optimized large language model developed by NVIDIA, derived from OpenAI's gpt-oss-120b
  • Hunyuan - a collection of Tencent's open-source efficient LLMs designed for versatile deployment across diverse computational environments
  • Phi-4 - a family of small language, multi-modal and reasoning models from Microsoft
  • NVIDIA 25 OpenReasoning-Nemotron - a collection of models from NVIDIA, trained on 5M reasoning traces for math, code and science
  • Kimi K2.5 - a collection of open-source, native multimodal agentic models from Moonshot AI that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration
  • GLM-5.2 - a Z.ai's flagship model for long-horizon tasks
  • Ling-3.0-flash - a native hybrid reasoning model from inclusionAI, operating with 124B total and 5.1B active parameters
  • Granite 4.1 - efficient language models from IBM for multilingual generation, coding, RAG, and AI assistant workflows
  • EXAONE-4.5 - LG's First Open-Weight Vision-Language Model for Industrial Intelligence
  • Step-3.5-Flash - most capable open-source foundation model, engineered to deliver frontier reasoning and agentic capabilities with exceptional efficiency
  • Nex-N2 - a collection of agent models built for real-world productivity scenarios

Back to Table of Contents

Coding

  • Qwen3-Coder-Next - a collection of Qwen's open-weight language models designed specifically for coding agents and local development
  • Mistral 20AI 23FA520F Devstral 2 - a couple of agentic LLMs for software engineering tasks, excelling at using tools to explore codebases, edit multiple files, and power SWE Agents
  • Mistral 20AI 23FA520F Mellum 2 - an assistant model trained by JetBrain
  • MiniMax-M3 - a native multimodal model with 1M context
  • MiniMax-M2 - a collection of SOTA models for real-world dev & agents
  • Laguna-S-2.1 - a 118B total parameter Mixture-of-Experts model with 8B activated parameters per token, designed for agentic coding and long-horizon work
  • SWE-FastContext - a family of code-search models from Microsoft powering the Explore subagent for coding agents
  • OmniCoder-9B - a 9-billion parameter coding agent model built by Tesslate, fine-tuned on top of Qwen3.5-9B's hybrid architecture
  • NousCoder-14B - a competitive programming model post-trained on Qwen3-14B via reinforcement learning
  • MusaCoder-27B - a code model developed by Moore Threads for PyTorch-to-CUDA/MUSA native kernel generation

Back to Table of Contents

Multimodal

  • Qwen3-Omni - a collection of the natively end-to-end multilingual omni-modal foundation models from Qwen
  • GLM-4.6V - a collection of open source multimodal models with native tool use from Zhipu AI

Back to Table of Contents

Image

  • Qwen-Image - a collection of models for image generation, edit and decomposition from Qwen
  • Qwen3-VL - a collection of the most powerful vision-language models in the Qwen series to date
  • GLM-Image - an image generation model
  • Granite Vision - multimodal models from IBM built for visual document analysis and image understanding
  • HunyuanImage - a collection of image generation models from Tencent
  • HunyuanVideo - a collection of video generation models from Tencent
  • Vidi - a collection of models for multimodal video understanding and creation
  • FastVLM - a collection of VLMs with efficient vision encoding from Apple
  • MiniCPM-o & MiniCPM-V - multimodal models with leading performance
  • LFM2-VL - a colection of vision-language models, designed for on-device deployment
  • ClipTagger-12b - a vision-language model (VLM) designed for video understanding at massive scale

Back to Table of Contents

Audio

  • OpenAI 23412991 whisper-large-v3 - a state-of-the-art model for automatic speech recognition (ASR) and speech translation from OpenAI
  • NVIDIA 25 Nemotron Speech - a collection of open, state-of-the-art, production‑ready enterprise speech models from NVIDIA for ASR, TTS, Speaker Diarization and S2SOpenAI
  • NVIDIA 25 NVIDIA NemotronLabs VoiceChat 11B - a 11B end-to-end, real-time speech full duplex (FD) model from NVIDIA for conversational AI that jointly performs streaming speech understanding and speech generation
  • Qwen3-ASR - a collection of models that support language identification and ASR for 52 languages and dialects
  • Qwen3-TTS - a collection of TTS models that cover 10 major languages as well as multiple dialectal voice profiles to meet global application needs
  • Granite Speech - a collection of compact and efficient speech-language models from IBM, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST)
  • Mistral 20AI 23FA520F Voxtral-Small-24B-2507 - an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance
  • Mistral 20AI 23FA520F Voxtral-Mini-4B-Realtime-2602 - a multilingual, realtime speech-transcription model and among the first open-source solutions to achieve accuracy comparable to offline systems with a delay of <500ms
  • Mistral 20AI 23FA520F Voxtral-4B-TTS-2603 - frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents
  • chatterbox - first production-grade open-source TTS model
  • VibeVoice - a collection of frontier text-to-speech models from Microsoft
  • Kitten TTS - a collection of open-source realistic text-to-speech models designed for lightweight deployment and high-quality voice synthesis
  • NVIDIA 25 Streaming Sortformer Diarizer 4spk v2.1 - a streaming version of a novel end-to-end neural model for speaker diarization from NVIDIA

Back to Table of Contents

Retrieval-Augmented Generation

  • NVIDIA 25 Nemotron RAG - a set of tools to build retrieval-augmented generation (RAG) systems, improve search and ranking accuracy, and extract structured data from complex docs
  • Qwen3-Embedding - a collection of the latest proprietary Qwen models, specifically designed for text embedding and ranking tasks
  • Qwen3-VL-Embedding - an addition to the Qwen embedding models, specifically designed for multimodal information retrieval and cross-modal understanding
  • Qwen3-Reranker - a collection of the latest proprietary Qwen models, engineered to refine embedding results
  • Qwen3-VL-Reranker - an addition to the Qwen embedding models, specifically designed for multimodal information retrieval and cross-modal understanding

Back to Table of Contents

Safeguards

  • Mistral 20AI 23FA520F Shieldstral 1.0 3B - a compact 3B-parameter, policy-adaptive multimodal safety classifier
  • Granite Guardian - a collection of safety models from IBM for detecting risks, toxicity, and hallucinations in LLM workflows
  • Qwen3Guard - a collection of safety moderation models built upon Qwen3
  • NVIDIA 25 NemoGuard - a collection of models from NVIDIA for content safety, topic-following and security guardrails
  • NVIDIA 25 Nemotron-3.5-Content-Safety - a small language model (SLM) that uses Google's Gemma-3-4B-it as the base and is fine-tuned by NVIDIA on multimodal, multilingual, and reasoning-oriented content-safety datasets
  • NVIDIA 25 Privasis - a collection of lightweight text-sanitization models from NVIDIA designed to remove or abstract sensitive information from text according to a user-provided sanitization instruction
  • SingGuard - a collection of policy-adaptive multimodal LLM Guardrails with dynamic reasoning
  • HARC - a family of safety-aligned instruction models from Microsoft trained with HARC
  • OpenAI 23412991 gpt-oss-safeguard - a collection of safety reasoning models built-upon gpt-oss from OpenAI
  • OpenAI 23412991 privacy-filter - a bidirectional token-classification model from OpenAI for personally identifiable information (PII) detection and masking in text
  • AprielGuard - a safeguard model designed to detect and mitigate both safety risks and security threats in LLM interactions

Back to Table of Contents

Miscellaneous

  • Intern-S2 - a collection of multimodal foundation models for scientific intelligence and long-horizon agents
  • Fara1.5 - a collection of multimodal computer use agents (CUA) for web browsers from Microsoft
  • Marco-MoE - a suit of multilingual MoE models with highly-sparse architectures
  • Jan-v3 - a 4B baseline model for fine-tuning, designed for downstream work: improved instruction following out of the box, strong starting point for fine-tuning and effective lightweight coding assistance
  • Jan-v2-VL - a family of VLM focused on reliable, many-step task execution
  • NVIDIA 25 Nemotron-Orchestrator-8B - a state-of-the-art 8B orchestration model designed to solve complex, multi-turn agentic tasks by coordinating a diverse set of expert models and tools
  • Arch-Router-1.5B - the fastest LLM router model that aligns to subjective usage preferences
  • Waypoint - a collection of real-time interactive video world models
  • Hunyuan3D - a collection of everything related (models, datasets etc.) to 3D assets generation from Tencent
  • Hunyuan-GameCraft-1.0 - a novel framework for high-dynamic interactive video generation in game environments
  • void-model - a model from Netflix that removes objects from videos along with all interactions they induce on the scene — not just secondary effects like shadows and reflections, but physical interactions like objects falling when a person is removed

Back to Table of Contents

Tools

Models

  • llmfit llmfit - hundreds of models & providers, one command to find what runs on your hardware
  • outlines outlines - structured outputs for LLMs
  • llama swap llama-swap - reliable model swapping for any local OpenAI compatible server - llama.cpp, vllm, etc.
  • llguidance llguidance - super-fast structured outputs

Back to Table of Contents

Agent Frameworks

  • AutoGPT AutoGPT - a powerful platform that allows you to create, deploy, and manage continuous AI agents that automate complex workflows
  • langflow langflow - a powerful tool for building and deploying AI-powered agents and workflows
  • langchain langchain - build context-aware reasoning applications
  • anything llm anything-llm - the all-in-one Desktop & Docker AI application with built-in RAG, AI agents, No-code agent builder, MCP compatibility, and more
  • autogen autogen - a programming framework for agentic AI
  • Flowise Flowise - build AI agents, visually
  • pi pi - AI agent toolkit: coding agent CLI, unified LLM API, TUI & web UI libraries, Slack bot, vLLM pods
  • llama index llama_index - the leading framework for building LLM-powered agents over your data
  • crewAI crewAI - a framework for orchestrating role-playing, autonomous AI agents
  • agno agno - a full-stack framework for building Multi-Agent Systems with memory, knowledge and reasoning
  • sim sim - open-source platform to build and deploy AI agent workflows
  • OpenAI 23412991 openai agents python openai-agents-python - a lightweight, powerful framework for multi-agent workflows
  • NVIDIA 25 NemoClaw NemoClaw - run OpenClaw more securely inside NVIDIA OpenShell with managed inference
  • SuperAGI SuperAGI - an open-source framework to build, manage and run useful Autonomous AI Agents
  • camel camel - the first and the best multi-agent framework
  • pydantic ai pydantic-ai - a Python agent framework designed to help you quickly, confidently, and painlessly build production grade applications and workflows with Generative AI
  • txtai txtai - all-in-one open-source AI framework for semantic search, LLM orchestration and language model workflows
  • agent framework agent-framework - a framework for building, orchestrating and deploying AI agents and multi-agent workflows with support for Python and .NET
  • archgw archgw - a high-performance proxy server that handles the low-level work in building agents: like applying guardrails, routing prompts to the right agent, and unifying access to LLMs, etc.
  • Google 234285F4 genkit genkit - open-source framework for building AI-powered apps in JavaScript, Go, and Python, built and used in production by Google
  • ClaraVerse ClaraVerse - privacy-first, fully local AI workspace with Ollama LLM chat, tool calling, agent builder, Stable Diffusion, and embedded n8n-style automation
  • NVIDIA 25 NeMo Agent Toolkit NeMo-Agent-Toolkit - an open-source library for efficiently connecting and optimizing teams of AI agents
  • ragbits ragbits - building blocks for rapid development of GenAI applications

Back to Table of Contents

Model Context Protocol

  • mindsdb mindsdb - federated query engine for AI - the only MCP Server you'll ever need
  • github mcp server github-mcp-server - GitHub's official MCP Server
  • playwright mcp playwright-mcp - Playwright MCP server
  • chrome devtools mcp chrome-devtools-mcp - Chrome DevTools for coding agents
  • n8n mcp n8n-mcp - a MCP for Claude Desktop / Claude Code / Windsurf / Cursor to build n8n workflows for you
  • mcp awslabs/mcp - AWS MCP Servers — helping you get the most out of AWS, wherever you use MCP
  • mcp atlassian mcp-atlassian - MCP server for Atlassian tools (Confluence, Jira)
  • dbhub dbhub - zero-dependency, token-efficient database MCP server for Postgres, MySQL, SQL Server, MariaDB, SQLite

Back to Table of Contents

Retrieval-Augmented Generation

  • pathway pathway - Python ETL framework for stream processing, real-time analytics, LLM pipelines and RAG
  • graphrag graphrag - a modular graph-based RAG system
  • LightRAG LightRAG - simple and fast RAG
  • haystack haystack - AI orchestration framework to build customizable, production-ready LLM applications, best suited for building RAG, question answering, semantic search or conversational agent chatbots
  • vanna vanna - an open-source Python RAG framework for SQL generation and related functionality
  • graphiti graphiti - build real-time knowledge graphs for AI Agents
  • onyx onyx - the AI platform connected to your company's docs, apps, and people
  • claude context claude-context - make entire codebase the context for any coding agent
  • pipeshub ai pipeshub-ai - a fully extensible and explainable workplace AI platform for enterprise search and workflow automation

Back to Table of Contents

Coding Agents

  • opencode opencode - a AI coding agent built for the terminal
  • zed zed - a next-generation code editor designed for high-performance collaboration with humans and AI
  • OpenHands OpenHands - a platform for software development agents powered by AI
  • cline cline - autonomous coding agent right in your IDE, capable of creating/editing files, executing commands, using the browser, and more with your permission every step of the way
  • aider aider - AI pair programming in your terminal
  • tabby tabby - an open-source GitHub Copilot alternative, set up your own LLM-powered code completion server
  • continue continue - create, share, and use custom AI code assistants with our open-source IDE extensions and hub of models, rules, prompts, docs, and other building blocks
  • void void - an open-source Cursor alternative, use AI agents on your codebase, checkpoint and visualize changes, and bring any model or host locally
  • goose goose - an open-source, extensible AI agent that goes beyond code suggestions
  • Roo Code Roo-Code - a whole dev team of AI agents in your code editor
  • crush crush - the glamourous AI coding agent for your favourite terminal
  • kilocode kilocode - open source AI coding assistant for planning, building, and fixing code
  • humanlayer humanlayer - the best way to get AI coding agents to solve hard problems in complex codebases
  • 99 99 - neovim AI agent done right
  • ProxyAI ProxyAI - the leading open-source AI copilot for JetBrains

Back to Table of Contents

Computer Use

  • open interpreter open-interpreter - a natural language interface for computers
  • OmniParser OmniParser - a simple screen parsing tool towards pure vision based GUI agent
  • openwork openwork - an open-source alternative to Claude Cowork, powered by OpenCode
  • cua cua - the Docker Container for Computer-Use AI Agents
  • Agent S Agent-S - an open agentic framework that uses computers like a human
  • self operating computer self-operating-computer - a framework to enable multimodal models to operate a computer
  • OpenRoom OpenRoom - a browser-based desktop where AI Agent operates every app through natural language, from MiniMaxAI

Back to Table of Contents

Browser Automation

  • puppeteer puppeteer - a JavaScript API for Chrome and Firefox
  • playwright playwright - a framework for Web Testing and Automation
  • browser use browser-use - make websites accessible for AI agents
  • firecrawl firecrawl - turn entire websites into LLM-ready markdown or structured data
  • stagehand stagehand - the AI Browser Automation Framework
  • nanobrowser nanobrowser - open-source Chrome extension for AI-powered web automation

Back to Table of Contents

Memory Management

  • mem0 mem0 - universal memory layer for AI Agents
  • mempalace mempalace - the highest-scoring AI memory system ever benchmarked
  • letta letta - the stateful agents framework with memory, reasoning, and context management
  • supermemory supermemory - memory engine and app that is extremely fast, scalable
  • cognee cognee - memory for AI Agents in 5 lines of code
  • LMCache LMCache - supercharge your LLM with the fastest KV Cache Layer
  • memU memU - an open-source memory framework for AI companions
  • Google 234285F4 reasoning bank reasoning-bank - a memory mechanism for agents that learns from both successful and failed trajectories, with reasoning stored as memory content

Back to Table of Contents

Testing, Evaluation and Observability

  • langfuse langfuse - an open-source LLM engineering platform: LLM Observability, metrics, evals, prompt management, playground, datasets. Integrates with OpenTelemetry, Langchain, OpenAI SDK, LiteLLM, and more
  • opik opik - debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards
  • openllmetry openllmetry - an open-source observability for your LLM application, based on OpenTelemetry
  • giskard giskard - an open-source evaluation & testing for AI & LLM systems
  • agenta agenta - an open-source LLMOps platform: prompt playground, prompt management, LLM evaluation, and LLM observability all in one place
  • NVIDIA 25 evaluator Evaluator - open-source library for scalable, reproducible evaluation of AI models and benchmarks

Back to Table of Contents

Research

  • Perplexica Perplexica - an open-source alternative to Perplexity AI, the AI-powered search engine
  • gpt researcher gpt-researcher - an LLM based autonomous agent that conducts deep local and web research on any topic and generates a long report with citations
  • SurfSense SurfSense - an open-source alternative to NotebookLM / Perplexity / Glean
  • open notebook open-notebook - an open-source implementation of Notebook LM with more flexibility and features
  • RD Agent RD-Agent - automate the most critical and valuable aspects of the industrial R&D process
  • local deep researcher local-deep-researcher - fully local web research and report writing assistant
  • local deep research local-deep-research - an AI-powered research assistant for deep, iterative research
  • maestro maestro - an AI-powered research application designed to streamline complex research tasks

Back to Table of Contents

Training and Fine-tuning

  • heretic heretic - fully automatic censorship removal for language models
  • sentence transformers sentence-transformers - a Python library for using and training embedding and reranker models for applications like retrieval augmented generation, semantic search, and more
  • trl trl - train transformer language models with reinforcement learning
  • OpenRLHF OpenRLHF - an easy-to-use, high-performance open-source RLHF framework built on Ray, vLLM, ZeRO-3 and HuggingFace Transformers, designed to make RLHF training simple and accessible
  • slime slime - an LLM post-training framework for RL Scaling
  • kiln Kiln - the easiest tool for fine-tuning LLM models, synthetic data generation, and collaborating on datasets
  • OpenEnv OpenEnv - an interface library for RL post training with environments
  • augmentoolkit augmentoolkit - train an open-source LLM on new facts
  • NVIDIA 25 rl RL - scalable toolkit for efficient model reinforcement
  • miles miles - an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime
  • NVIDIA 25 gym Gym - evaluate and improve models and agents using environments
  • SpecForge SpecForge - train speculative decoding models effortlessly and port them smoothly to SGLang serving

Back to Table of Contents

Security and Sandboxing

  • NVIDIA 25 garak garak - the LLM vulnerability scanner from NVIDIA
  • NVIDIA 25 Guardrails Guardrails - an open-source toolkit from NVIDIA for easily adding programmable guardrails to LLM-based conversational systems
  • NVIDIA 25 OpenShell OpenShell - the safe, private runtime for autonomous AI agents from NVIDIA
  • CubeSandbox CubeSandbox - instant, concurrent, secure & lightweight sandbox for AI agents

Back to Table of Contents

Miscellaneous

  • context7 context7 - up-to-date code documentation for LLMs and AI code editors
  • deepwiki open deepwiki-open - open source DeepWiki: AI-powered wiki generator for GitHub/Gitlab/Bitbucket repositories
  • cai cai - Cybersecurity AI (CAI), the framework for AI Security
  • speakr speakr - a personal, self-hosted web application designed for transcribing audio recordings
  • presenton presenton - an open-source AI presentation generator and API
  • OmniGen2 OmniGen2 - exploration to advanced multimodal generation
  • 4o ghibli at home 4o-ghibli-at-home - a powerful, self-hosted AI photo stylizer built for performance and privacy
  • Observer Observer - local open-source micro-agents that observe, log and react, all while keeping your data private and secure
  • mobile use mobile-use - a powerful, open-source AI agent that controls your Android or IOS device using natural language
  • gabber gabber - build AI applications that can see, hear, and speak using your screens, microphones, and cameras as inputs
  • promptcat promptcat - a zero-dependency prompt manager/catalog/library in a single HTML file

Back to Table of Contents

Hardware

  • UCajiMK CY9icRhLepS8 3ug Alex Ziskind - tests of pcs, laptops, gpus etc. capable of running LLMs
  • UCiaQzXI5528Il6r2NNkrkJA Digital Spaceport - reviews of various builds designed for LLM inference
  • UCP0QFok6EimQYTMj5qOLNow Donato Capitella - practical and insightful tutorials on running LLMs locally
  • UCQs0lwV6E4p7LQaGJ6fgy5Q JetsonHacks - information about developing on NVIDIA Jetson Development Kits
  • UC8h2Sf yyo1WXeEUr OHgyg Miyconst - tests of various types of hardware capable of running LLMs
  • Kolosal - LLM Memory calculator - estimate the RAM requirements of any GGUF model instantly
  • LLM Inference VRAM & GPU Requirement Calculator - calculate how many GPUs you need to deploy LLMs
  • ZLUDA ZLUDA - CUDA on non-NVIDIA GPUs
  • Strix Halo AI Toolboxes - toolboxes for GenAI on AMD Ryzen AI MAX+: containerized environments for LLMs, Image Generation, and Fine-tuning
  • Strix Halo Wiki - a website to gather important information and practical guides for systems powered by AMD Ryzen AI MAX and MAX+ processors
  • ai notes ai-notes - random AI notes for working with local models or playing around with random machine learning bits

Back to Table of Contents

Tutorials

Models

Back to Table of Contents

Prompt Engineering

Back to Table of Contents

Context Engineering

  • Context Engineering Context-Engineering - a frontier, first-principles handbook inspired by Karpathy and 3Blue1Brown for moving beyond prompt engineering to the wider discipline of context design, orchestration, and optimization
  • Awesome Context Engineering Awesome-Context-Engineering - a comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems

Back to Table of Contents

Inference

  • production stack vLLM Production Stack - vLLM’s reference system for K8S-native cluster-wide deployment with community-driven performance optimization

Back to Table of Contents

Agents

  • superpowers superpowers - an agentic skills framework & software development methodology that works
  • GenAI Agents GenAI Agents - tutorials and implementations for various Generative AI Agent techniques
  • 500 AI Agents Projects 500+ AI Agent Projects - a curated collection of AI agent use cases across various industries
  • 12 factor agents 12-Factor Agents - principles for building reliable LLM applications
  • agents towards production Agents towards production - end-to-end, code-first tutorials covering every layer of production-grade GenAI agents, guiding you from spark to scale with proven patterns and reusable blueprints for real-world launches
  • agents agents.md - a simple, open format for guiding coding agents
  • agentskills Agent Skills - a simple, open format for giving agents new capabilities and expertise
  • skills skills - Hugging Face Skills are definitions for AI/ML tasks like dataset creation, model training and evaluation
  • LLM Agents Ecosystem Handbook LLM Agents & Ecosystem Handbook - one-stop handbook for building, deploying, and understanding LLM agents with 60+ skeletons, tutorials, ecosystem guides, and evaluation tools
  • Google 234285F4 601 real-world gen AI use cases - 601 real-world gen AI use cases from the world's leading organizations by Google
  • OpenAI 23412991 A practical guide to building agents - a practical guide to building agents by OpenAI

Back to Table of Contents

Retrieval-Augmented Generation

  • llm app Pathway AI Pipelines - ready-to-run cloud templates for RAG, AI pipelines, and enterprise search with live data
  • RAG Techniques RAG Techniques - various advanced techniques for Retrieval-Augmented Generation (RAG) systems
  • Controllable RAG Agent Controllable RAG Agent - an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering that uses sophisticated graph based algorithm to handle the tasks
  • langchain rag cookbook LangChain RAG Cookbook - a collection of modular RAG techniques, implemented in LangChain + Python

Back to Table of Contents

Miscellaneous

Back to Table of Contents

Communities

Back to Table of Contents

Contributing

We welcome contributions! Please see CONTRIBUTING.md for guidelines on how to get started.