只需修改一行代码,即可将 GPT 替换为任何其他大语言模型。Xinference 允许您在云端、本地或笔记本电脑上运行开源模型、语音模型以及多模态模型——所有操作均通过一个统一且可用于生产环境的推理 API 完成。
repos=xorbitsai/inference&type=Date)](https://star-history.com/#xorbitsai/inference&Date)
Mistral 模型的官方推理库
``` ### Local ``` cd $HOME && git clone https://github.com/mistralai/mistral-inference cd $HOME/mistral-inference && poetry install . ```
一款用于大语言模型的高吞吐、低内存消耗的推理与服务引擎
. --- ## About vLLM is a fast and easy-to-use library for LLM inference and serving.
将任何电脑或边缘设备转变为您的计算机视觉项目的指挥中心。
These licenses are listed in [`inference/models`](/inference/models).
OpenAI Whisper 模型的 C/C++ 版本端口
only ## Real-time audio input example This is a naive example of performing real-time inference on audio from your microphone.
大型语言模型文本生成推理
/pkgs/container/text-generation-inference) - [AMD](https://github.com/huggingface/text-generation-inference/pkgs/container/text-generation-inference
DeepSpeed 是一个深度学习优化库,它能够让分布式训练与推理变得更加简单、高效且富有成效。
: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.
让大型人工智能模型的开发成本更低、速度更快,且更易被大众使用
">Inference</a> <ul> <li><a href="#Colossal-Inference">Colossal-Inference: Large AI Models Inference Speed Doubled</a></li> <li
基于 PyTorch 的 Ultralytics YOLOv5,用于目标检测、实例分割、分类、模型训练以及结果导出。
# Run inference on a text file listing stream URLs python detect.py --weights yolov5s.pt --source list.streams # Run inference using a glob
SGLang 是一种用于大型语言模型和多模态模型的高性能服务框架。
/)). - [2026/02] 🔥 Unlocking 25x Inference Performance with SGLang on NVIDIA GB300 NVL72 ([blog](https://lmsys.org/blog/2026-02-20-gb300-inferencex
借助 CTranslate2,实现更快速的 Whisper 文本转录
faster-whisper** is a reimplementation of OpenAI's Whisper model using [CTranslate2](https://github.com/OpenNMT/CTranslate2/), which is a fast inference
Triton Inference Server 提供了一种优化的云端及边缘推理解决方案。
# Triton Inference Server Triton Inference Server is an open source inference serving software that streamlines AI inferencing.
机器学习工程开放教材
Inference** 1. **[Inference](inference)** - model inference insights **Part 6. Development** 1.
MLPerf® 推理基准测试的参考实现
# MLPerf® Inference Benchmark Suite MLPerf Inference is a benchmark suite for measuring how fast systems can run models in a variety of deployment
🎨 专为 TypeScript 设计的强大模式匹配库,具备智能类型推断功能。
align="center"> The exhaustive Pattern Matching library for <a href="https://github.com/microsoft/TypeScript">TypeScript</a> with smart type inference
轻松地对 Qwen、Gemma 或任何开源权重大型语言模型进行微调、评估与部署!
/vision/qwen2_5_vl_3b/inference/vllm_infer.yaml) • [Inference](configs/recipes/vision/qwen2_5_vl_3b/inference/infer.yaml) | | Qwen2-VL 2B | [