ITADN
找到 1.0 万 条结果
X

只需修改一行代码,即可将 GPT 替换为任何其他大语言模型。Xinference 允许您在云端、本地或笔记本电脑上运行开源模型、语音模型以及多模态模型——所有操作均通过一个统一且可用于生产环境的推理 API 完成。

repos=xorbitsai/inference&type=Date)](https://star-history.com/#xorbitsai/inference&Date)

8176 751活跃于 2025-10-08
M
mistralai/mistral-inferenceJupyter Notebook已归档

Mistral 模型的官方推理库

``` ### Local ``` cd $HOME && git clone https://github.com/mistralai/mistral-inference cd $HOME/mistral-inference && poetry install . ```

8939 874活跃于 2025-10-08
V

一款用于大语言模型的高吞吐、低内存消耗的推理与服务引擎

. --- ## About vLLM is a fast and easy-to-use library for LLM inference and serving.

5.7 万 1.1 万活跃于 2026-08-04
F
1861 436活跃于 2025-10-08
R

将任何电脑或边缘设备转变为您的计算机视觉项目的指挥中心。

These licenses are listed in [`inference/models`](/inference/models).

1829 200活跃于 2026-08-03
G

OpenAI Whisper 模型的 C/C++ 版本端口

only ## Real-time audio input example This is a naive example of performing real-time inference on audio from your microphone.

3.6 万 4340活跃于 2025-10-09
H

大型语言模型文本生成推理

/pkgs/container/text-generation-inference) - [AMD](https://github.com/huggingface/text-generation-inference/pkgs/container/text-generation-inference

9702 1213活跃于 2025-10-08
D

DeepSpeed 是一个深度学习优化库,它能够让分布式训练与推理变得更加简单、高效且富有成效。

: Enabling Efficient Inference of Transformer Models at Unprecedented Scale.

3.1 万 3727活跃于 2026-08-03
H

让大型人工智能模型的开发成本更低、速度更快,且更易被大众使用

">Inference</a> <ul> <li><a href="#Colossal-Inference">Colossal-Inference: Large AI Models Inference Speed Doubled</a></li> <li

2.6 万 3115活跃于 2025-10-08
U

基于 PyTorch 的 Ultralytics YOLOv5,用于目标检测、实例分割、分类、模型训练以及结果导出。

# Run inference on a text file listing stream URLs python detect.py --weights yolov5s.pt --source list.streams # Run inference using a glob

2.2 万 6073活跃于 2025-10-08
S

SGLang 是一种用于大型语言模型和多模态模型的高性能服务框架。

/)). - [2026/02] 🔥 Unlocking 25x Inference Performance with SGLang on NVIDIA GB300 NVL72 ([blog](https://lmsys.org/blog/2026-02-20-gb300-inferencex

1.9 万 3178活跃于 2026-08-04
S

借助 CTranslate2,实现更快速的 Whisper 文本转录

faster-whisper** is a reimplementation of OpenAI's Whisper model using [CTranslate2](https://github.com/OpenNMT/CTranslate2/), which is a fast inference

1.8 万 1509活跃于 2025-10-08
I
798 16活跃于 2025-09-30
T

Triton Inference Server 提供了一种优化的云端及边缘推理解决方案。

# Triton Inference Server Triton Inference Server is an open source inference serving software that streamlines AI inferencing.

5008 657活跃于 2026-08-03
S

机器学习工程开放教材

Inference** 1. **[Inference](inference)** - model inference insights **Part 6. Development** 1.

1.4 万 822活跃于 2026-08-03
G

适用于实时媒体和流媒体领域的跨平台、可定制的机器学习解决方案。

1.2 万 1698活跃于 2025-10-09
H
3816 323活跃于 2025-10-08
M

MLPerf® 推理基准测试的参考实现

# MLPerf® Inference Benchmark Suite MLPerf Inference is a benchmark suite for measuring how fast systems can run models in a variety of deployment

653 196活跃于 2025-10-08
G

🎨 专为 TypeScript 设计的强大模式匹配库,具备智能类型推断功能。

align="center"> The exhaustive Pattern Matching library for <a href="https://github.com/microsoft/TypeScript">TypeScript</a> with smart type inference

8274 109活跃于 2025-10-08
O

轻松地对 Qwen、Gemma 或任何开源权重大型语言模型进行微调、评估与部署!

/vision/qwen2_5_vl_3b/inference/vllm_infer.yaml) • [Inference](configs/recipes/vision/qwen2_5_vl_3b/inference/infer.yaml) | | Qwen2-VL 2B | [

8203 647活跃于 2026-08-04