Petals: Run LLMs at home, BitTorrent-style

·Hacker News··

Petals Run large language models at home, BitTorrent‑style Generate text with Llama 3.1 (up to 405B), Mixtral (8x22B), Falcon (40B+) or BLOOM (176B) and fine‑tune them for your tasks — using a consumer-grade GPU or Google Colab. You load a part of the model, then join a network of people serving its other parts. Single‑batch inference runs at up to 6 tokens/sec for Llama 2 (70B) and up to 4 tokens/sec for Falcon (180B) — enough for chatbots and interact

Read full article →

Related Articles

LG to ban residential proxies from smart TV apps
DemiGuru · Hacker News · 1d ago
Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
piotrgrabowski · Hacker News · 1d ago
Apple defeats liability for not scanning iCloud for CSAM
speckx · Hacker News · 1d ago
Everyone should know SIMD
WadeGrimridge · Hacker News · 12h ago
GigaToken: ~1000x faster Language model tokenization
syrusakbary · Hacker News · 13h ago