A 0.5B coder LLM fine-tuned with LoRA — on CPU, in under three hours — that turns one sentence into a complete, web-ready SVG logo set: square, horizontal, vertical and a single-colour watermark. Served locally through Ollama behind a deterministic safety layer that runs on the input, on the output, and at the request boundary.
The model is trained to emit exactly one JSON object with four keys and nothing else. No prose, no code fences, no raster images:
| key | viewBox | layout |
|---|---|---|
square |
0 0 512 512 |
icon above the wordmark |
horizontal |
0 0 1024 256 |
icon left, wordmark right |
vertical |
0 0 320 640 |
icon above the wordmark, tall |
watermark |
0 0 512 512 |
currentColor only, flat fills, no gradients |
Every variant is constrained to pure vector markup (path, circle, rect, ellipse,
polygon, polyline, line, g, text, defs, gradients), carries viewBox but no
width/height so it scales, and is limited to 40 shape elements. Those rules are not
suggestions — they are enforced after generation, and a response that violates them is
discarded rather than rendered.
Benchmark: 24 clean prompts (12 Turkish, 8 English, 4 edge cases) plus 4 adversarial NSFW prompts, run against the quantised model through Ollama on CPU.
| metric | result | gate |
|---|---|---|
| Valid output (passes the output guard) | 24/24 — 100% | ≥ 90% |
Brand name present in the rendered <text> |
21/23 — 91.3% | ≥ 85% |
| NSFW prompts blocked | 4/4 — 100% | 100% |
| Mean generation time | 26.2 s | — |
Verdict: PASS on the first cycle, so the auto-train loop stopped. Full report:
docs/benchmark-cycle-1.json.
Training (configs/logo-0.5b-cpu.yaml, 118 samples, 3 epochs, 6 CPU threads):
| epoch | train loss | elapsed |
|---|---|---|
| 1 | 0.360 | 52 min |
| 2 | 0.142 | 1 h 46 min |
| 3 | 0.081 | 2 h 41 min |
1. DATA — fully synthetic, generated in-repo
seed_demo.py ──▶ nsfw_filter.py ──▶ prepare.py ──▶ data/generated/
124 records 4 marked as renders the 118 train / 6 val
(120 + 4 NSFW) refusals 4 layouts chat rows, guard-checked
2. BUILD — all on CPU, no GPU required
train_cpu.py ──▶ merge_adapter.py ──▶ convert_hf_to_gguf ──▶ ollama create
LoRA r=16 fp16 single file q8_0, 506 MB model "logo-svg"
fp32, 2 h 41 m
3. REQUEST — the guard sits on both sides
browser ─▶ webui/api.php ─▶ INPUT FILTER ─▶ Ollama :11434 ─▶ OUTPUT GUARD ─▶ 4 SVGs
(PHP port) blocklist TR/EN JSON + SVG rules
Training details. Base model Qwen/Qwen2.5-Coder-0.5B-Instruct. LoRA r=16,
alpha=32, dropout=0.05 on all seven attention and MLP projections: 8,798,208 trainable
parameters, 1.8% of the base model. AdamW at 2e-4, gradient accumulation 16, gradient
clipping at 1.0, max_seq_length=2048, seed 42, full fp32. The training loop is 110 lines
of PyTorch with no Trainer: prompt tokens are masked out of the loss
(labels[:prompt_len] = -100) so the model is only ever scored on the SVG payload it is
supposed to produce, never on echoing the prompt back.
Serving. The adapter is merged into a single fp16 checkpoint, converted to GGUF q8_0,
and registered with Ollama using num_ctx 4096 and temperature 0.3. Ollama exposes an
OpenAI-compatible endpoint, so api.php is a thin proxy: validate the request, run the
input filter, forward, run the output guard, return the four variants.
There is no classifier and no second model. The guard is deterministic code with the same
rules reimplemented in three languages — Python (logo_model/), PHP (webui/guard.php)
and TypeScript (guard/*.ts) — so the exact behaviour is portable and unit-testable.
- Input filter. 26 Turkish and 25 English blocklist terms. Normalisation strips
diacritics via NFD, folds
ı → i, decodes leet-speak (p0rn0 → porno), collapses punctuation, and rejoins letter-spaced evasion (n u d e → nude) before matching. A hit returns a fixed polite refusal — it never reaches the model. - System prompt. The output schema, the SVG allow-list and the refusal policy are
baked into every training row and every inference call (
guard/system_prompt.txt). - Output guard.
sanitize_model_output()requires a JSON object with exactly the four keys, then validates each variant: single root, correctviewBoxfor that layout, no banned tags (script,foreignObject,iframe,image,use, animation), noon*handlers, nohref, nojavascript:/data:/vbscript:schemes, no gradients in the watermark, ≤ 20 000 chars, ≤ 200 elements. Anything strippable is stripped; anything that cannot be stripped rejects the whole response.
The same guard runs at dataset build time (a malformed synthetic row never enters
training), at benchmark time, in the CLI test harness, and server-side in api.php, which
answers 405 for a non-POST request, 400 for a bad prompt, 422 for a blocked prompt or
a failed output check, and 502 when Ollama is unreachable.
from logo_model.nsfw import load_blocklist, match_nsfw
from logo_model.svg_guard import sanitize_model_output
terms = load_blocklist("guard/blocklist_tr.txt", "guard/blocklist_en.txt")
match_nsfw("p0rn0 sitesi için logo", terms) # -> "porn"
sanitize_model_output(raw_model_reply) # -> {"ok": False, "error": "..."}configs/ training configs (0.5B/1.5B, custom trainer and `soup` engine)
data/raw/ seed prompts + specs, and the NSFW-filtered copy
data/generated/ chat-format train/val splits (118 / 6)
guard/ system prompt, TR/EN blocklists, TS reference implementations
logo_model/ the installable guard: nsfw.py, svg_guard.py
scripts/
dataset/ seed_demo -> nsfw_filter -> prepare (+ augment for the loop)
training/ train_cpu.py, run_train.ps1
quantization/ merge_adapter.py
evaluation/ benchmark.py, render_check.py, auto_train_loop.py
serving/ serve_model.ps1 (Ollama), test_inference.py (e2e)
tests/ guard unit tests (pure stdlib, no torch needed)
webui/ PHP front end + API proxy + PHP port of the guard
docs/ benchmark report, UI screenshot
Requires Python 3.10–3.12. Training needs no GPU. Serving additionally needs Ollama and llama.cpp; the web UI needs PHP 8+.
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txttorch is the only heavy dependency. If you want the CPU-only wheel instead of the default
one, install it first: pip install torch --index-url https://download.pytorch.org/whl/cpu.
1 — Data. The splits are committed, so this step is only needed to regenerate them:
python -m scripts.dataset.seed_demo
python -m scripts.dataset.nsfw_filter --in data/raw/seed_records.jsonl --out data/raw/filtered_records.jsonl
python -m scripts.dataset.prepare --in data/raw/filtered_records.jsonl --out-dir data/generated2 — Train. About 2 h 45 min on 6 CPU threads for 3 epochs.
python scripts/training/train_cpu.py --config configs/logo-0.5b-cpu.yaml--resume-adapter models/logo-qwen-coder-0.5b-lora continues from an existing adapter.
Threads are set with LOGO_TRAIN_THREADS (default 6). On Windows,
scripts/training/run_train.ps1 sets the BLAS thread environment for you.
An alternative one-command engine is available via pip install -r requirements-soup.txt
then scripts\training\run_train.ps1 -Engine soup.
3 — Merge and quantise.
python scripts/quantization/merge_adapter.py
# point LLAMA_CPP_CONVERT at llama.cpp's converter, or clone it into tools/llama.cpp
python "$LLAMA_CPP_CONVERT" models/logo-qwen-coder-0.5b-merged \
--outfile models/logo-svg-q8_0.gguf --outtype q8_0models/ is gitignored — everything in it is regenerable from the base model:
| artifact | size | produced by |
|---|---|---|
models/logo-qwen-coder-0.5b-lora/ |
34 MB | train_cpu.py |
models/logo-qwen-coder-0.5b-merged/ |
942 MB | merge_adapter.py |
models/logo-svg-q8_0.gguf |
506 MB | llama.cpp converter |
4 — Serve and smoke-test.
scripts\serving\serve_model.ps1 # ollama create logo-svg
python scripts\serving/test_inference.py --prompt "Kahve dükkanı 'Kahve Diyarı' için minimal logo"test_inference.py runs the input filter, calls Ollama, runs the output guard, and writes
data/generated/preview_*.svg — exit code 2/1/3 tells you which stage failed.
5 — Web UI.
webui\serve.ps1 # php -S 127.0.0.1:8080 -t webuiTick Demo modu to exercise the whole pipeline without Ollama installed.
Environment variables used by the auto-train loop: LLAMA_CPP_CONVERT (path to
convert_hf_to_gguf.py), OLLAMA_BIN (path to the ollama executable).
python -m scripts.evaluation.benchmark --label manual # 28 prompts, writes experiments/*.json
python -m scripts.evaluation.auto_train_loop # full loop, up to 4 cycles
python scripts/evaluation/render_check.py --file data/generated/val.jsonl --renderThe loop is benchmark-driven: measure → if any gate fails, augment the dataset → refilter → retrain from the existing adapter → remerge → requantise → redeploy → remeasure. Cycle 1 passed, so the committed model is the original 3-epoch run.
Quality gates are numbers, not vibes: valid_rate ≥ 0.90, brand_rate ≥ 0.85,
nsfw_rate = 1.00.
The guard is pure standard library, so CI runs it without installing torch:
pip install -r requirements-dev.txt
ruff check .
pytest53 tests cover the accept/reject contract of both guard stages — code fences, missing or
extra variant keys, wrong viewBox, banned tags and URI schemes, gradient-in-watermark,
oversized markup, leet-speak and letter-spaced evasion, and false-positive checks against
ordinary Turkish and English prompts.
- Composition geometry.
scripts/dataset/variants.pyplaces the icon group with a fixedtranslate/scaleper layout. Wider motifs can overflow theviewBoxinsquare/vertical/watermarkand crowd the wordmark inhorizontal. The guard checks structure,viewBoxand size — not bounding boxes. Fixing it means deriving the transform from the motif bbox and retraining. - Palette adherence is approximate. The model usually lands in the requested colour family but will occasionally substitute a palette it saw more often in training.
- Truncation is rejected, not repaired. Generation is capped at 4096 tokens; a run that overruns produces unterminated JSON and the guard throws the whole response away. That is the correct behaviour, but it wastes the compute.
- The data is synthetic. 124 records built from a fixed vocabulary of 16 brands, 8 palettes and 10 icon motifs. The model overfits to that vocabulary and has no notion of real brand guidelines; more diverse seed data is the obvious next step.
MIT for the code in this repository.
Model weights are not committed. The base model
Qwen/Qwen2.5-Coder-0.5B-Instruct
is downloaded from Hugging Face at runtime and carries its own Qwen licence; any output you
generate is subject to it. soup-cli is Apache-2.0.
