Repository navigation

#

ggml

ggml-org/llama.cpp

LLM inference in C/C++

C++
87194
11 小时前

Replace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the cloud, on-premises, or even on your laptop.

Python
8595
4 天前

[Unmaintained, see README] An ecosystem of Rust libraries for working with large language models

Rust
6137
1 年前

llama and other large language models on iOS and MacOS offline using GGML library.

C
1884
16 天前

INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model

C++
1548
6 个月前

Calculate token/s & GPU memory requirement for any LLM. Supports llama.cpp/ggml/bnb/QLoRA quantization

JavaScript
1367
10 个月前

Suno AI's Bark model in C/C++ for fast text-to-speech generation

C++
842
1 年前

Whisper Dart is a cross platform library for dart and flutter that allows converting audio to text / speech to text / inference from Open AI models

C++
606
8 个月前

Run inference on MPT-30B using CPU

Python
577
2 年前

Port of MiniGPT4 in C++ (4bit, 5bit, 6bit, 8bit, 16bit CPU inference with GGML)

C++
569
2 年前

This custom_node for ComfyUI adds one-click "Virtual VRAM" for any UNet and CLIP loader as well MultiGPU integration in WanVideoWrapper, managing the offload/Block Swap of layers to DRAM *or* VRAM to maximize the latent space of your card. Also includes nodes for directly loading entire components (UNet, CLIP, VAE) onto the device you choose

Python
544
19 小时前

Running any GGUF SLMs/LLMs locally, on-device in Android

Kotlin
526
15 天前

CLIP inference in plain C/C++ with no extra dependencies

C++
523
4 个月前

Large Language Models for All, 🦙 Cult and More, Stay in touch !

HTML
445
2 年前

WIP Library Text To Speech From Suno AI's Bark in C/C++ for fast inference

C++
371
1 年前