Local inference
Ollama ↗
A local runtime for downloading, serving and testing a broad set of open models.
Open Source Radar
A reviewed watchlist for repositories that matter to the open-model ecosystem. Star counts and repository update times make it easy to compare the projects readers are watching.
Local inference
A local runtime for downloading, serving and testing a broad set of open models.
Model framework
Hugging Face’s model-definition library for training and inference across text, vision and audio.
AI interface
A self-hosted interface for working with Ollama and OpenAI-compatible model APIs.
Agent framework
A widely used platform for building applications and agent workflows around language models.
Generative media
A node-based graphical interface and backend for modular diffusion-model workflows.
Local inference
A compact C/C++ implementation for efficient local LLM inference across many devices.
Inference serving
A high-throughput, memory-efficient serving engine for large language and multimodal models.
RAG & agents
A retrieval-augmented generation engine that combines document retrieval with agent capabilities.
Fine-tuning
A unified toolkit for fine-tuning and evaluating a broad range of language and vision-language models.
Agent framework
Microsoft’s framework for programming multi-agent applications and tool-driven workflows.
RAG & data
A framework for document agents, retrieval workflows and connecting LLMs to data.
Vector database
A cloud-native vector database for scalable approximate-nearest-neighbour search.
Training & inference
A distributed deep-learning optimization library for efficient training and inference.
Observability
An open AI-engineering platform for LLM evaluation, observability and prompt management.
Inference serving
A high-performance serving framework for large language and multimodal models.
Vector database
Open search infrastructure for building AI applications around embeddings and retrieval.
Agent framework
Microsoft’s SDK for integrating LLMs, memory and tool use into applications.
Apple silicon
Apple’s array framework for machine learning research and efficient on-device model work.
RAG & agents
A modular orchestration framework for production-ready retrieval and agent pipelines.
Fine-tuning
Hugging Face’s parameter-efficient fine-tuning library for adapting foundation models.
Inference serving
NVIDIA’s optimized runtime and APIs for efficient LLM inference on its GPUs.
Training & inference
A collection of high-performance LLM recipes for pretraining, fine-tuning and deployment.
Inference serving
A server for exposing open-source models through OpenAI-compatible cloud APIs.
Inference serving
Hugging Face’s production text-generation server for deploying large language models.
Inference serving
NVIDIA’s optimized cloud and edge serving platform for a range of model frameworks.
Evaluation
An LLM evaluation platform spanning models and more than one hundred datasets.