TESIGN
SAVED

EN · KO

Auto-generated — not yet edited · only the numbers are verified

  1. TESIGN / RADAR
  2. REPOSITORY CARD
  3. makazhanalpamys/soup

makazhanalpamys/soup

Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.

★ 6.3k▲ +82 / 7d▲ 139 HNDaily star increases over 14 days. Hollow bars mean no data. The last day may be partial.2026-09-03 +122026-09-04 +402026-09-05 +472026-09-06 +382026-09-07 +32026-09-08 no data2026-09-09 +32026-09-10 +82026-09-11 +162026-09-12 +232026-09-13 +102026-09-14 no data2026-09-15 +232026-09-16 +2largest one-day increase in the 14 days+47last 14 days · stars per dayhollow = no data

RANKS Rising 30d #90

AT A GLANCE

LANGUAGE
Python
LICENSE
Apache-2.0
USAGE
Commercial use, redistribution OK. Keep notices; mark changes.
ACTIVITY
last commit 3 days ago ()
TOPICS
  • cli
  • consumer-gpu
  • dpo
  • fine-tuning
  • gguf
  • huggingface
  • llm
  • llmops
  • local-ai
  • local-llm
  • lora
  • low-vram
HOMEPAGE
https://trysoup.dev
REPOSITORY
GitHub ↗
CATEGORY
AI

A summary, not legal advice.

README EXCERPT

🌍 English Türkçe Soup Fine-tune and post-train LLMs in one command. No SSH, no config hell. Website · Quick Start · Web UI · Config · Docs · Commands · Models · Discord · Telegram · Product Hunt --- Soup turns the pain of LLM fine-tuning into a simple workflow. One config, one command, done. Fine-tune an 8B model on a 4 GB laptop GPU. Layer streaming keeps the frozen base out of VRAM and feeds it to the GPU one decoder layer at a time. Measured on an RTX 3050 Laptop 4 GB: Llama-3.1-8B-Instruct + NF4 at 119.6 tok/s, 3.32 GB peak — bit-exact against a normal resident run, and reproduced independently on an H100 at 113.00 tok/s in the same 3.32 GB. (The tok/s figure was measured on v0.72.2, before the v0.73.0 correctness repair that cost −4.8% at 32B; it has not been re-run on a 4 GB card since.) Opt-in ( stream layers: true ) and still BETA — how it works · all measurements · paper · check it yourself on a free Colab T4 (caps the process to 4 GB, then asserts a streamed model is bit-identical to a normal one) Llama-3.1-8B-Instruct + NF4, LoRA, batch 1, seq 512 on an RTX 3050 Laptop 4 GB — 3.32 GB peak, 119.6 tok/s . Full…

The opening of the GitHub README as stored, at most 1,200 characters. Markdown is not rendered.

TIMELINE

  1. Repository created
  2. FIRST SEEN BY TESIGN ◌ BACK CATALOG
  3. Last push

AI chooses the lists under the owner’s delegation. No human review is running in September 2026. Total stars, increases, cross-source signals, last updates and licences are shown as evidence. How ranks work →

If this repository gets editorial text (why, build, who, start, caveat) it becomes an edited entry. Until then the page shows only stored GitHub metadata and numbers.