| Kornia | Library is composed by a subset of packages containing operators that can be inserted within neural networks to train models to perform image transformations, epipolar geometry, depth estimation, and low-level image processing such as filtering and edge detection that operate directly on tensors | | |  | 05.08.2026 |
| Retrieval based Voice Conversion WebUI | An easy-to-use Voice Conversion framework based on VITS | 源文雨 | |  | 04.08.2026 |
| SGLang | System comprising a frontend language and runtime for efficiently programming and executing complex structured language model programs with optimizations for cache reuse and structured output decoding. | others |    , , , , , , , , , , , , ,  ,    ,  - website
 - 1, 2, 3, 4, 5, 6
|  | 04.08.2026 |
| SHAP | SHapley Additive exPlanations is a game theoretic approach to explain the output of any machine learning model | | |  | 03.08.2026 |
| YOLOv8 | State-of-the-art model that builds upon the success of previous YOLO versions and introduces new features and improvements to further boost performance and flexibility | Glenn Jocher | |  | 01.08.2026 |
| YOLOv3 | You Only Look Once | Glenn Jocher | |  | 29.07.2026 |
| YOLOv5 | You Only Look Once | Glenn Jocher | |  | 29.07.2026 |
| dm_control | DeepMind Infrastructure for Physics-Based Simulation | others | |  | 28.07.2026 |
| ADK | Collection provides ready-to-use agents built on top of the Agent Development Kit, designed to accelerate your development process | | |  | 24.07.2026 |
| OmegaConf | Hierarchical configuration system, with support for merging configurations from multiple sources providing a consistent API regardless of how the configuration was created | Omry Yadan | |  | 22.07.2026 |
| Agent Starter Pack | Collection of production-ready Generative AI Agent templates built for Google Cloud | Kristopher Overholt | |  | 21.07.2026 |
| AlphaFold | Highly accurate protein structure prediction | others | |  | 14.07.2026 |
| GigaAM | SSL pretraining framework that leverages masked language modeling with targets derived from a speech recognition model | | |  | 14.07.2026 |
| s | This paper introduces O-Voxel, a new sparse voxel representation and compression framework that enables high-fidelity, efficient 3D asset generation with flexible geometry and detailed appearance from learned compact latent spaces | others | |  | 09.07.2026 |
| NeMo | A conversational AI toolkit built for researchers working on automatic speech recognition, natural language processing, and text-to-speech synthesis | others | |  | 07.07.2026 |
| Anime Face Detector | Anime Face Detector using mmdet and mmpose | hysts | |  | 05.07.2026 |
| Chronos-2 | Pretrained, zero-shot time series forecasting model that uses group attention and synthetic multivariate training to perform univariate, multivariate, and covariate-informed forecasting with state-of-the-art accuracy across diverse real-world benchmarks. | others | |  | 02.07.2026 |
| Nano Banana | An image generation and editing model powered by generative artificial intelligence and developed by Google DeepMind | Guillaume Vernade | |  | 30.06.2026 |
| Presidio | Context aware, pluggable and customizable PII de-identification service for text and images | Omri Mendels | , ,  ,  , , ,  , , , ,  - 1, 2
|  | 28.06.2026 |
| Weaver | Lightweight autoregressive adapter for factorized draft models that constructs proposal trees from top-K marginals to restore conditional dependencies and enable faster speculative decoding with optimized tree verification and CUDA kernels in SGLang. | | , , , ,   , , , , , , , , , , , , , ,  ,   ,  ,  - website
 ,  - 1, 2, 3, 4, 5, 6, 7, 8, 9
|  | 24.06.2026 |
| TransformerLens | Library for doing mechanistic interpretability of GPT-2 Style language models | | |  | 22.06.2026 |
| Hyperopt | Python library for serial and parallel optimization over awkward search spaces, which may include real-valued, discrete, and conditional dimensions | | |  | 21.06.2026 |
| highway-env | A collection of environments for autonomous driving and tactical decision-making tasks | Edouard Leurent | |  | 19.06.2026 |
| SentencePiece | An unsupervised text tokenizer and detokenizer mainly for Neural Network-based text generation systems where the vocabulary size is predetermined prior to the neural model training | | |  | 14.06.2026 |
| RF-DETR | Lightweight, real-time detection transformer that uses weight-sharing neural architecture search to automatically discover optimal accuracy-latency tradeoffs for object detection across diverse target datasets. | | , , , , ,  - blog
 ,    , , , ,  - website
 - 1
|  | 13.06.2026 |
| Mem0 | Self-improving memory layer for LLM applications, enabling personalized AI experiences that save costs and delight users | | |  | 13.06.2026 |
| Swarm | Educational framework exploring ergonomic, lightweight multi-agent orchestration | others | |  | 07.06.2026 |
| Duo | This software project implements Duo, a diffusion-based language modeling framework that improves discrete diffusion text generation using Gaussian-guided curriculum learning and Discrete Consistency Distillation for faster training and few-step sampling. | others | |  | 03.06.2026 |
| Magenta RT | An open-weights live music model that allows you to interactively create, control and perform music in the moment | Chris Donahue | |  | 02.06.2026 |
| TorchGeo | PyTorch domain library that provides datasets, transforms, samplers, and pre-trained models specific to geospatial data | others | |  | 01.06.2026 |
| pymdp | Package for simulating Active Inference agents in Markov Decision Process environments | others | |  | 29.05.2026 |
| ActionMesh | Temporal 3D diffusion framework that generates production-ready, topology-consistent animated 3D meshes from inputs like video, text, or static 3D shapes in a fast, feed-forward manner. | |  , ,  , , , , ,  ,   - 1, 2, 3
|  | 28.05.2026 |
| Lyria 2 | Delivers high-fidelity music and professional-grade audio, capturing subtle nuances across a range of genres and intricate compositions | Katie Nguyen | |  | 13.05.2026 |
| Video Seal | Comprehensive framework for neural video watermarking and a competitive open-sourced model | | |  | 11.05.2026 |
| Google Cloud Text-to-Speech | Enables easy integration of Google text recognition technologies into developer applications | | |  | 07.05.2026 |
| Imagen 4 | Text-to-image model, with photorealistic images, near real-time speed, and sharper clarity | Katie Nguyen | |  | 06.05.2026 |
| CrewAI | Lean, lightning-fast Python framework built entirely from scratch—completely independent of LangChain or other agent frameworks | João Moura |   ,  ,  - website
, , , , , , , ,  - 1, 2, 3
|  | 20.04.2026 |
| CodeGemma | How to load, fine-tune and deploy CodeGemma model on SQL by utilising Hugging Face | Carlo Fisicaro | |  | 20.04.2026 |
| Hello, many worlds | This tutorial shows how a classical neural network can learn to correct qubit calibration errors | Michael Broughton | |  | 18.04.2026 |
| Text Generation Web UI | The best local UI for large language models, with easy setup and powerful features. 100% offline. | oobabooga |   , , , , , ,  , , ,   - 1, 2, 3, 4
|  | 13.04.2026 |
| DataChain | AI-dataframe to enrich, transform and analyze data from cloud storages for ML training and LLM apps | Daniel K | |  | 13.04.2026 |
| JAX MD | Differentiable physics and molecular dynamics simulation framework in JAX that enables scalable, GPU-accelerated simulations and end-to-end optimization of entire trajectories, with flexible primitives and neural network integration. | | , , , , , ,     , ,   - 1, 2, 3, 4
|  | 05.04.2026 |
| GraphCast | Learning skillful medium-range global weather forecasting | others |  , , , ,  , , , ,  - 1, 2, 3
|  | 30.03.2026 |
| ignite | High-level library to help with training and evaluating neural networks in PyTorch flexibly and transparently. | Anmol Joshi |    , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , , ,  ,     - 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11
|  | 25.03.2026 |
| edgeai-tensorlab | Edge AI Model Development Tools | Manu | |  | 25.03.2026 |
| Dopamine | Research framework for fast prototyping of reinforcement learning algorithms | | |  | 24.03.2026 |
| autoresearch | Bilevel Autoresearch is a bilevel autoresearch framework in which an outer autoresearch loop reads and modifies the code and traces of an inner autoresearch loop to inject new Python search mechanisms at runtime, thereby improving the inner loop’s task search behavior without changing the underlying LLM. | Andrej Karpathy |  , , , , , ,  ,  ,  , ,  - 1, 2, 3, 4, 5, 6
|  | 19.03.2026 |
| Opik | From RAG chatbots to code assistants to complex agentic pipelines and beyond, build LLM systems that run better, faster, and cheaper with tracing, evaluations, and dashboards | Jacques Verré | |  | 19.03.2026 |
| PEFT | Parameter-Efficient Fine-Tuning methods enable efficient adaptation of pre-trained language models to various downstream applications without fine-tuning all the model's parameters | | |  | 17.03.2026 |
| tada | Generative text-to-speech and spoken language modeling framework that introduces a one-to-one synchronized text-acoustic tokenization scheme for unified large language model-based speech modeling with reduced hallucinations and inference cost. | others | |  | 16.03.2026 |
| spiky | This project introduces a spiking neural network paradigm that reframes modern AI models in terms of spike-based polychronization to achieve combinatorially large encoding capacity and dramatically higher energy efficiency than conventional artificial neural networks. | Anatoli S tarostin | |  | 14.03.2026 |
| DeepFloyd IF | State-of-the-art open-source text-to-image model with a high degree of photorealism and language understanding | others | |  | 12.03.2026 |
| Diffusers | Provides pretrained diffusion models across multiple modalities, such as vision and audio, and serves as a modular toolbox for inference and training of diffusion models | Hugging Face | , , , ,  , , ,  , , , ,   - 1
|  | 12.03.2026 |
| CatBoost | High-performance open source library for gradient boosting on decision trees | others | |  | 09.03.2026 |
| BEiT | Self-supervised vision representation model, which stands for Bidirectional Encoder representation from Image Transformers | | |  | 07.03.2026 |
| Diffusion_models_tutorial | This project develops diffusion probabilistic models for high-quality image synthesis, leveraging a new connection to denoising score matching with Langevin dynamics to achieve state-of-the-art generative performance and a progressive lossy decompression scheme. | | |  | 05.03.2026 |
| HACRL-code | This paper presents Heterogeneous Agent Collaborative Reinforcement Learning, a Reinforcement Learning from Verifiable Reward framework in which heterogeneous agents share verified rollouts during training to collaboratively optimize while still executing independently at inference time. | others | |  | 05.03.2026 |
| ExecuTorch | PyTorch's unified solution for deploying AI models on-device, from smartphones to microcontrollers, built for privacy, performance, and portability | SaoirseARM | |  | 04.03.2026 |
| PaperBanana | Agentic framework that uses advanced vision-language and image-generation models to automatically create and refine publication-ready academic illustrations, evaluated on a new benchmark of methodology diagrams and statistical plots. | others |   , , , , , ,  ,  , , , , , ,  - 1, 2, 3, 4, 5, 6, 7
|  | 03.03.2026 |
| GigaAgent | Универсальный агент-оркестратор для решения широкого круга задач (ReAct + REPL) | Mikelarg | |  | 26.02.2026 |
| Giskard | Open-source library to detect hallucinations and security issues to turn them into test suites that you can automatically execute | | |  | 17.02.2026 |
| SGLang | Fast serving framework for large language models and vision language models | others | |  | 15.02.2026 |
| Thorsten-Voice | Free to use, offline working, high quality german TTS voice should be available for every project without any license struggling | Thorsten Müller |  , , , , ,  ,  - website
 , ,  - 1, 2, 3, 4
|  | 02.02.2026 |
| fastai | The fastai deep learning library | Sylvain Gugger | |  | 28.01.2026 |
| LM Evaluation Harness | Framework for few-shot evaluation of language models. | Lintang Sutawika | |  | 27.01.2026 |
| OpenSpiel | Collection of environments and algorithms for research in general reinforcement learning and search/planning in games | others | |  | 22.01.2026 |
| Keras | Multi-backend deep learning framework, with support for JAX, TensorFlow, PyTorch, and OpenVINO (for inference-only) | François Chollet | |  | 21.01.2026 |
| Grounding DINO | Marrying DINO with Grounded Pre-Training for Open-Set Object Detection | others |  , , , , , ,  , , ,  , , , 
|  | 19.01.2026 |
| GNN | Production-tested library for building GNNs at large scale | others | |  | 14.01.2026 |
| Gin Config | Lightweight configuration framework for Python, based on dependency injection | | |  | 14.01.2026 |
| PyGlove | General-purpose library for Python object manipulation | others | |  | 14.01.2026 |
| Brax | A differentiable physics engine that simulates environments made up of rigid bodies, joints, and actuators | others | |  | 14.01.2026 |
| T5 | Text-To-Text Transfer Transformer | others | |  | 14.01.2026 |
| SeqIO | Library for processing sequential data to be fed into downstream sequence models | others | |  | 14.01.2026 |
| TFDS | Collection of ready-to-use datasets for use with TensorFlow, Jax, and other Machine Learning frameworks | Ryan Sepassi | |  | 14.01.2026 |
| TensorFlow Privacy | Library that includes implementations of TensorFlow optimizers for training machine learning models with differential privacy | | |  | 14.01.2026 |
| TF-Agents | A reliable, scalable and easy to use TensorFlow library for Contextual Bandits and Reinforcement Learning | others | |  | 14.01.2026 |
| Reverb | Efficient and easy-to-use data storage and transport system designed for machine learning research | others | |  | 14.01.2026 |
| ACME | A library of reinforcement learning components and agents | others | |  | 14.01.2026 |
| Gemma | Family of open-weights Large Language Model by Google DeepMind, based on Gemini research and technology | Google |   , ,  , ,   , , , , , , ,  - 1, 2, 3, 4
|  | 14.01.2026 |
| Sonnet | Library built on top of TensorFlow 2 designed to provide simple, composable abstractions for machine learning research | others | |  | 14.01.2026 |
| T5X | Modular, composable, research-friendly framework for high-performance, configurable, self-service training, evaluation, and inference of sequence models at many scales | others | |  | 14.01.2026 |
| TFRS | Library for building recommender system models using TensorFlow | Maciej Kula | |  | 14.01.2026 |
| CycleGAN | This notebook demonstrates unpaired image to image translation using conditional GAN's | | |  | 14.01.2026 |
| Neural style transfer | This tutorial uses deep learning to compose one image in the style of another image | | |  | 14.01.2026 |
| Pix2Pix | This notebook demonstrates image to image translation using conditional GAN's | | |  | 14.01.2026 |
| XLA | Accelerated Linear Algebra is an open-source machine learning compiler for GPUs, CPUs, and ML accelerators | George Karpenkov | |  | 13.01.2026 |
| TF-DF | TensorFlow Decision Forests is a library to train, run and interpret decision forest models (e.g., Random Forests, Gradient Boosted Trees) in TensorFlow | | |  | 12.01.2026 |
| DDSP | Differentiable Digital Signal Processing library, which enables direct integration of classic signal processing elements with deep learning methods | | |  | 09.01.2026 |
| Langfun | PyGlove powered library that aims to make language models fun to work with | Daiyi Peng | |  | 09.01.2026 |
| circuit-tracer | Library implements tools for finding circuits using features from (cross-layer) MLP transcoders | |   ,   , , , , , , ,  - 1, 2, 3
|  | 08.01.2026 |
| MuJoCo | A general purpose physics engine that aims to facilitate research and development in robotics, biomechanics, graphics and animation, machine learning, and other areas which demand fast and accurate simulation of articulated structures interacting with their environment | | |  | 07.01.2026 |
| NNCF | Neural Network Compression Framework is a PyTorch-based toolkit that applies methods like sparsity, quantization, and binarization with fine-tuning to produce hardware-efficient neural network models that accelerate inference while preserving accuracy. | | |  | 07.01.2026 |
| Transformer Engine | Library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point precision on Hopper, Ada, and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference | | , ,   , , , , , , , , ,    - 1, 2, 3, 4
|  | 06.01.2026 |
| AlphaEvolve | Evolutionary coding agent that substantially enhances capabilities of state-of-the-art LLMs on highly challenging tasks such as tackling open scientific problems or optimizing critical pieces of computational infrastructure | others | |  | 05.01.2026 |
| OpenForecaster | Language-model-based system trained on automatically generated forecasting questions from news to improve the accuracy, calibration, and consistency of open-ended predictions about future events. | | |  | 04.01.2026 |
| SAM Audio | Foundation model for general audio source separation that integrates text, visual, and temporal prompts, enabling flexible and state-of-the-art separation of diverse sounds across multiple domains. | others | ,  ,  , , , , , ,  , ,  - 1, 2
|  | 30.12.2025 |
| TTUR | Two time-scale update rule for training GANs with stochastic gradient descent on arbitrary GAN loss functions | | |  | 14.12.2025 |
| SamGeo | A Python package for segmenting geospatial data with the Segment Anything Model (SAM) | Qiusheng Wu |  , , , , ,   , , , , ,  , ,  - 1, 2, 3
|  | 10.12.2025 |
| DiffSynth | Restructured architectures including Text Encoder, UNet, VAE, among others, maintaining compatibility with models from the open-source community while enhancing computational performance | | |  | 04.12.2025 |
| Segment Anything 3 | Unified model that detects, segments, and tracks objects in images and videos based on concept prompts, which we define as either short noun phrases (e.g., “yellow school bus”), image exemplars, or a combination of both | others | |  | 19.11.2025 |
| optax | This project introduces AlgoPerf, a competitive time-to-result benchmark designed to reliably compare and identify state-of-the-art neural network training algorithms across multiple workloads on fixed hardware. | others | |  | 14.11.2025 |
| PyTerrier | A Python framework for performing information retrieval experiments | | |  | 13.11.2025 |
| Transfer learning and fine-tuning | You will learn how to classify images of cats and dogs by using transfer learning from a pre-trained network | François Chollet | |  | 11.11.2025 |
| Datasets | A Community Library for Natural Language Processing | others | |  | 10.11.2025 |
| Vertex AI | Search brings together the power of deep information retrieval, state-of-the-art natural language processing, and the latest in LLM processing to understand user intent and return the most relevant results for the user | Megha Agarwal | |  | 06.11.2025 |
| AlphaEvolve problems | LLM-guided evolutionary coding system that autonomously discovers and optimizes mathematical constructions and solutions to complex open problems across multiple fields of mathematics. | | , , , ,  , ,   - 1, 2, 3, 4, 5, 6, 7, 8
|  | 05.11.2025 |
| TRL | Set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step, Reward Modeling step to the Proximal Policy Optimization step | others | |  | 05.11.2025 |
| SAE Lens | Training Sparse Autoencoders on Language Models | | |  | 27.10.2025 |
| CSBDeep | Toolbox for content-aware restoration of fluorescence microscopy images (CARE), based on deep learning via Keras and TensorFlow | others | |  | 25.10.2025 |
| HYPIR | Image restoration framework that initializes a restoration model from a pre-trained diffusion model and fine-tunes it with adversarial training to achieve fast, high-fidelity, and controllable image restoration in a single forward pass. | others | |  | 16.10.2025 |
| Intel® Neural Compressor | Aims to provide popular model compression techniques such as quantization, pruning (sparsity), distillation, and neural architecture search on mainstream frameworks such as TensorFlow, PyTorch, ONNX Runtime, and MXNet, as well as Intel extensions such as Intel Extension for TensorFlow and Intel Extension for PyTorch | intel | , ,   , , ,  ,     , , , , , 
|  | 10.10.2025 |
| Edgeai for beginners | This course is designed to guide beginners through the exciting world of Edge AI, covering fundamental concepts, popular models, inference techniques, device-specific applications, model optimization, and the development of intelligent Edge AI agents. | Lee Stott | ,  , , , , , , , , , , , ,  - 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17
|  | 08.10.2025 |
| LangChain | Framework for developing applications powered by large language models | Bagatur | , , ,  , ,     , , , , , , , ,  - 1
|  | 03.10.2025 |
| Compel | Text prompt weighting and blending library for transformers-type text embedding systems | Damian Stewart | |  | 02.10.2025 |
| OWL-ViT | Simple Open-Vocabulary Object Detection with Vision Transformers | others | |  | 29.09.2025 |
| PaddleSpeech | Open-source toolkit on PaddlePaddle platform for a variety of critical tasks in speech and audio, with the state-of-art and influential models | others | |  | 28.09.2025 |
| LlamaIndex | Data framework for your LLM application | Jerry Liu | |  | 26.09.2025 |
| CWM | Code World Model, a 32-billion-parameter open-weights LLM, to advance research on code generation with world models | others | |  | 24.09.2025 |
| Qwen3-VL | Family of advanced vision-language models with long-context, interleaved multimodal support (text, images, video) and improved architectures that deliver state-of-the-art multimodal understanding and reasoning across diverse benchmarks and real-world applications. | others | , , , ,     , , , , ,  , , ,  , , , , , , , ,  - website
- 1, 2, 3, 4, 5, 6, 7, 8
|  | 23.09.2025 |
| SAHI | A lightweight vision library for performing large scale object detection & instance segmentation | others | |  | 23.09.2025 |
| Qwen3-Omni | Single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relative to single-modal counterparts | others | |  | 22.09.2025 |
| WMAR | Custom tokenizer-detokenizer finetuning procedure that improves RCC, and a complementary watermark synchronization layer | | |  | 19.09.2025 |
| LIMIT | On the Theoretical Limitations of Embedding-Based Retrieval | | |  | 27.08.2025 |
| fast-stable-diffusion | fast-stable-diffusion + DreamBooth | Ben | |  | 25.08.2025 |
| TensorRT | SDK for high-performance deep learning inference, includes a deep learning inference optimizer and runtime that delivers low latency and high throughput for inference applications | nvidia | |  | 19.08.2025 |
| DINOv3 | Produces high-quality dense features that achieve outstanding performance on various vision tasks, significantly surpassing previous self- and weakly-supervised foundation models | others | |  | 14.08.2025 |
| A2A | Google's Agent-to-Agent protocol, a standardized way for AI agents to communicate and collaborate | | , , , , , ,    , , , , , , , ,  - 1, 2
|  | 07.08.2025 |
| pytorch-lightning | Pretrain, finetune ANY AI model of ANY size on 1 or 10,000+ GPUs with zero code changes. | Aniket Maurya | ,   , , , ,  , , , , , , , ,     ,  - website
 ,  - 1, 2, 3
|  | 04.08.2025 |
| clean-fid | Investigates how low-level image processing choices—particularly aliased resizing and lossy compression—significantly and unpredictably affect GAN evaluation metrics like FID, and provides signal-processing-based recommendations and reference code for more reliable generative model assessment. | | , , , , ,  , , , , ,  - project
 , 
|  | 02.08.2025 |
| MinerU | Open-source solution for high-precision document content extraction | others |    , , , , , , , , , ,  ,   - website
|  | 31.07.2025 |
| Llama 4 | Open-weight natively multimodal models with unprecedented context length support and our first built using a MoE architecture | meta | |  | 29.07.2025 |
| VC | Client software for performing real-time voice conversion using various Voice Conversion AI | w-okada | |  | 19.07.2025 |
| Rembg | Tool to remove image backgrounds | Arhenniuss | |  | 16.07.2025 |
| guidance | Enables you to control modern language models more effectively and efficiently than traditional prompting or chaining | Scott Lundberg | |  | 16.07.2025 |
| Hogwild! Inference | Run LLM "workers" in parallel, allowing them to synchronize via a concurrently-updated attention cache and prompt these workers to decide how best to collaborate | others | |  | 15.07.2025 |
| Autodistill | Uses big, slower foundation models to train small, faster supervised models | autodistill |  , , , , , , , , , , , , , , , ,  , ,  - 1
|  | 10.07.2025 |
| ComfyUI | Powerful and modular stable diffusion GUI and backend | comfyanonymous | |  | 09.07.2025 |
| AutoGen | Framework that enables development of LLM applications using multiple agents that can converse with each other to solve tasks | microsoft | |  | 07.07.2025 |
| Image classification | This tutorial shows how to classify images of flowers | Billy Lamberta | |  | 01.07.2025 |
| Hunyuan | Open-source large language model built on a fine-grained Mixture-of-Experts architecture | manayang | |  | 01.07.2025 |
| RL Games | High performance RL library | | |  | 26.06.2025 |
| Whisper | Automatic speech recognition system trained on 680,000 hours of multilingual and multitask supervised data collected from the web | others | |  | 26.06.2025 |
| IT³ | Idempotent Test-Time Training, approach that enables on-the-fly adaptation to distribution shifts using only the current test instance, without any auxiliary task design | others | |  | 25.06.2025 |
| V-JEPA 2 | Self-supervised approach that combines internet-scale video data with a small amount of interaction data, to develop models capable of understanding, predicting, and planning in the physical world | FAIR | |  | 11.06.2025 |
| Crawl4AI | LLM Friendly Web Crawler & Scrapper | UncleCode | |  | 10.06.2025 |
| LightAutoML | Allows you create machine learning models using just a few lines of code, or build your own custom pipeline using ready blocks | |   ,  , , , , , ,    - website
, , , , 
|  | 06.06.2025 |
| BEIR | Heterogeneous benchmark that evaluates the zero-shot out-of-distribution generalization of diverse information retrieval models across 18 text retrieval datasets and multiple retrieval paradigms. | | ,  ,  , ,   - 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13
|  | 04.06.2025 |
| Eso-LMs | Family of diffusion-based language models that fuse autoregressive and masked diffusion paradigms using causal attention to enable exact likelihood computation, KV caching, and parallel generation for improved speed-quality trade-offs in unconditional text generation. | others | |  | 03.06.2025 |
| MDLM | The project implements simple masked discrete diffusion language models using an effective training recipe and a simplified Rao-Blackwellized objective to train encoder-only models that generate text efficiently and achieve state-of-the-art diffusion-based language modeling performance approaching autoregressive perplexity. | others | |  | 27.05.2025 |
| Gorilla | Finetuned LLaMA-based model that surpasses the performance of GPT-4 on writing API calls | | |  | 26.05.2025 |
| Amphion | Amphion is an open-source, beginner-friendly toolkit that provides a unified, extensible framework for audio, music, and speech generation, supporting tasks like text-to-speech, text-to-audio, and singing voice conversion with pretrained models and essential processing components. | others | , , , , , , , , , , , , , , , , , , , , , , , , , ,   , , , , , , , , ,  , , , ,  - project
- website
 - 1, 2, 3
|  | 26.05.2025 |
| TimesFM | Time-series foundation model for forecasting whose out-of-the-box zero-shot performance on a variety of public datasets comes close to the accuracy of state-of-the-art supervised forecasting models for each individual dataset | | |  | 26.05.2025 |
| LLaMA Factory | Easy-to-use and efficient platform for training and fine-tuning large language models | others |    , , , , , , ,  ,     , , , , , 
|  | 19.05.2025 |
| Llama 3.1 | First openly available model that rivals the top AI models when it comes to state-of-the-art capabilities in general knowledge, steerability, math, tool use, and multilingual translation | unsloth | |  | 17.05.2025 |
| Phi-3.5 | 3.8 billion parameter language model trained on 3.3 trillion tokens, whose overall performance, as measured by both academic benchmarks and internal testing, rivals that of models such as Mixtral 8x7B and GPT-3.5, despite being small enough to be deployed on a phone | unsloth | |  | 16.05.2025 |
| DPO Zephyr | Starting from a dataset of outputs ranked by a teacher model, we apply distilled direct preference optimization to learn a chat model with significantly improved intent alignment | others | |  | 16.05.2025 |
| Ray | Unified framework for scaling AI and Python applications | others | |  | 16.05.2025 |
| cutlass | CUDA Templates and Python DSLs for High-Performance Linear Algebra | Kihiro Bando | , ,  , , ,  - website
- 1, 2, 3, 4, 5, 6, 7, 8
|  | 13.05.2025 |
| LangGraph | Library for building stateful, multi-actor applications with LLMs, used to create agent and multi-agent workflows | William FH | |  | 09.05.2025 |
| CircleGuardBench | First-of-its-kind benchmark for evaluating the protection capabilities of large language model guard systems | White Circle | |  | 06.05.2025 |
| Optimum | Extension of Transformers and Diffusers, providing a set of optimization tools enabling maximum efficiency to train and run models on targeted hardware, while keeping things easy to use | Hugging Face | |  | 29.04.2025 |
| Qwen2.5-Omni | End-to-end multimodal model designed to perceive diverse modalities, including text, images, audio, and video, while simultaneously generating text and natural speech responses in a streaming manner | others | |  | 29.04.2025 |
| verl | HybridFlow, which combines single-controller and multi-controller paradigms in a hybrid manner to enable flexible representation and efficient execution of the RLHF dataflow | others | , , ,   , , , , , , ,   
|  | 29.04.2025 |
| Composable-Diffusion | Compositional generation framework that treats diffusion models as energy-based components which can be combined to generate complex, photorealistic scenes with precise object relations and attribute bindings beyond those seen during training. | | |  | 24.04.2025 |
| EAT | Emotional Adaptation for Audio-driven Talking-head method, which transforms emotion-agnostic talking-head models into emotion-controllable ones in a cost-effective and efficient manner through parameter-efficient adaptations | | |  | 22.04.2025 |
| Trackers | clean, modular re-implementations of leading multi-object tracking algorithms released under the permissive Apache 2.0 license | Piotr Skalski | |  | 16.04.2025 |
| Evidently | An open-source framework to evaluate, test and monitor ML models in production | | |  | 08.04.2025 |
| Generative AI Knowledge Base | How to extract question & answer pairs out of documents using Generative AI | David Cavazos | |  | 04.04.2025 |
| Applied AI Engineering | Reference guides, blueprints, code samples, and hands-on labs developed by the Google Cloud Applied AI Engineering team | Mikhail Chertushkin | |  | 01.04.2025 |
| Moshi | Speech-text foundation model and full-duplex spoken dialogue framework | others | , , , ,  - demo
, , ,     , , , , 
|  | 31.03.2025 |
| xFormers | Toolbox to Accelerate Research on Transformers | others | |  | 25.03.2025 |
| BiRefNet | Bilateral reference framework for high-resolution dichotomous image segmentation | others | |  | 24.03.2025 |
| ESM | Evolutionary Scale Modeling: Pretrained language models for proteins | others | |  | 21.03.2025 |
| SkyThought | Train your own O1 preview model within $450 | Sumanth Hegde | |  | 20.03.2025 |
| SigLIP 2 | Family of new multilingual vision-language encoders that build on the success of the original SigLIP | others | |  | 17.03.2025 |
| Sentence Transformers | Multilingual Sentence, Paragraph, and Image Embeddings using BERT & Co | | |  | 07.03.2025 |
| DeepLabCut | Efficient method for markerless pose estimation based on transfer learning with deep neural networks that achieves excellent results with minimal training data | others | |  | 28.02.2025 |
| TPOT | Automated Machine Learning tool that optimizes machine learning pipelines using genetic programming | | |  | 24.02.2025 |
| piper | A fast, local neural text to speech system | Mateo Cedillo | |  | 24.02.2025 |
| Multimodal Maestro | Gives you more control over large multimodal models to get the outputs you want | Piotr Skalski | |  | 18.02.2025 |
| YuE | Groundbreaking series of open-source foundation models designed for music generation, specifically for transforming lyrics into full songs | Mozer | |  | 16.02.2025 |
| YOLOv12 | Attention-centric, real-time object detection framework that matches the speed of CNN-based YOLO models while significantly improving accuracy across multiple model scales. | | |  | 06.02.2025 |
| moondream | Tiny vision language model that kicks ass and runs anywhere | Vik Korrapati | |  | 06.02.2025 |
| Building Your Own Federated Learning Algorithm | We discuss how to implement federated learning algorithms without deferring to the tff.learning API | Zachary Charles | |  | 29.01.2025 |
| Custom Federated Algorithms, Part 1: Introduction to the Federated Core | This tutorial is the first part of a two-part series that demonstrates how to implement custom types of federated algorithms in TensorFlow Federated using the Federated Core - a set of lower-level interfaces that serve as a foundation upon which we have implemented the Federated Learning layer | Krzysztof Ostrowski | |  | 29.01.2025 |
| Custom Federated Algorithms, Part 2: Implementing Federated Averaging | This tutorial is the second part of a two-part series that demonstrates how to implement custom types of federated algorithms in TFF using the Federated Core, which serves as a foundation for the Federated Learning layer | Krzysztof Ostrowski | |  | 29.01.2025 |
| Federated Learning for Image Classification | We use the classic MNIST training example to introduce the Federated Learning API layer of TFF, tff.learning - a set of higher-level interfaces that can be used to perform common types of federated learning tasks, such as federated training, against user-supplied models implemented in TensorFlow | Krzysztof Ostrowski | |  | 29.01.2025 |
| Federated Learning for Text Generation | We start with a RNN that generates ASCII characters, and refine it via federated learning | Krzysztof Ostrowski | |  | 29.01.2025 |
| STAR | Spatial Temporal Augmentation with T2V models for Real-world video super-resolution, a novel approach that leverages T2V models for real-world video super-resolution, achieving realistic spatial details and robust temporal consistency | others | |  | 22.01.2025 |
| InvSR | Image super-resolution technique based on diffusion inversion, aiming at harnessing the rich image priors encapsulated in large pre-trained diffusion models to improve SR performance | | |  | 21.01.2025 |
| Invariant Agent Stack | Framework-less approach that currently consists of three key projects, each of which can be used independently or in combination to build, test, and secure AI agents | Fei Xie | |  | 20.01.2025 |
| Contextualized Topic Models | Family of topic models that use pre-trained representations of language to support topic modeling | | |  | 16.01.2025 |
| Ollama | Get up and running with large language models | Michael Yang | |  | 13.01.2025 |
| NotebookLlama | Open Source version of NotebookLM | Sanyam Bhutani | |  | 09.01.2025 |
| Anomalib | Deep learning library that aims to collect state-of-the-art anomaly detection algorithms for benchmarking on both public and private datasets | others | |  | 08.01.2025 |