Skip to content

Repository files navigation

logo-model

CI license python

A 0.5B coder LLM fine-tuned with LoRA — on CPU, in under three hours — that turns one sentence into a complete, web-ready SVG logo set: square, horizontal, vertical and a single-colour watermark. Served locally through Ollama behind a deterministic safety layer that runs on the input, on the output, and at the request boundary.

Logo Atölyesi web UI


Output contract

The model is trained to emit exactly one JSON object with four keys and nothing else. No prose, no code fences, no raster images:

key viewBox layout
square 0 0 512 512 icon above the wordmark
horizontal 0 0 1024 256 icon left, wordmark right
vertical 0 0 320 640 icon above the wordmark, tall
watermark 0 0 512 512 currentColor only, flat fills, no gradients

Every variant is constrained to pure vector markup (path, circle, rect, ellipse, polygon, polyline, line, g, text, defs, gradients), carries viewBox but no width/height so it scales, and is limited to 40 shape elements. Those rules are not suggestions — they are enforced after generation, and a response that violates them is discarded rather than rendered.

Results

Benchmark: 24 clean prompts (12 Turkish, 8 English, 4 edge cases) plus 4 adversarial NSFW prompts, run against the quantised model through Ollama on CPU.

metric result gate
Valid output (passes the output guard) 24/24 — 100% ≥ 90%
Brand name present in the rendered <text> 21/23 — 91.3% ≥ 85%
NSFW prompts blocked 4/4 — 100% 100%
Mean generation time 26.2 s —

Verdict: PASS on the first cycle, so the auto-train loop stopped. Full report: docs/benchmark-cycle-1.json.

Training (configs/logo-0.5b-cpu.yaml, 118 samples, 3 epochs, 6 CPU threads):

epoch train loss elapsed
1 0.360 52 min
2 0.142 1 h 46 min
3 0.081 2 h 41 min

How it works

  1. DATA — fully synthetic, generated in-repo
     seed_demo.py ──▶ nsfw_filter.py ──▶ prepare.py ──▶ data/generated/
     124 records       4 marked as        renders the      118 train / 6 val
     (120 + 4 NSFW)    refusals           4 layouts        chat rows, guard-checked

  2. BUILD — all on CPU, no GPU required
     train_cpu.py ──▶ merge_adapter.py ──▶ convert_hf_to_gguf ──▶ ollama create
     LoRA r=16        fp16 single file      q8_0, 506 MB         model "logo-svg"
     fp32, 2 h 41 m

  3. REQUEST — the guard sits on both sides
     browser ─▶ webui/api.php ─▶ INPUT FILTER ─▶ Ollama :11434 ─▶ OUTPUT GUARD ─▶ 4 SVGs
                 (PHP port)       blocklist TR/EN                  JSON + SVG rules

Training details. Base model Qwen/Qwen2.5-Coder-0.5B-Instruct. LoRA r=16, alpha=32, dropout=0.05 on all seven attention and MLP projections: 8,798,208 trainable parameters, 1.8% of the base model. AdamW at 2e-4, gradient accumulation 16, gradient clipping at 1.0, max_seq_length=2048, seed 42, full fp32. The training loop is 110 lines of PyTorch with no Trainer: prompt tokens are masked out of the loss (labels[:prompt_len] = -100) so the model is only ever scored on the SVG payload it is supposed to produce, never on echoing the prompt back.

Serving. The adapter is merged into a single fp16 checkpoint, converted to GGUF q8_0, and registered with Ollama using num_ctx 4096 and temperature 0.3. Ollama exposes an OpenAI-compatible endpoint, so api.php is a thin proxy: validate the request, run the input filter, forward, run the output guard, return the four variants.

Safety layer

There is no classifier and no second model. The guard is deterministic code with the same rules reimplemented in three languages — Python (logo_model/), PHP (webui/guard.php) and TypeScript (guard/*.ts) — so the exact behaviour is portable and unit-testable.

  1. Input filter. 26 Turkish and 25 English blocklist terms. Normalisation strips diacritics via NFD, folds ı → i, decodes leet-speak (p0rn0 → porno), collapses punctuation, and rejoins letter-spaced evasion (n u d e → nude) before matching. A hit returns a fixed polite refusal — it never reaches the model.
  2. System prompt. The output schema, the SVG allow-list and the refusal policy are baked into every training row and every inference call (guard/system_prompt.txt).
  3. Output guard. sanitize_model_output() requires a JSON object with exactly the four keys, then validates each variant: single root, correct viewBox for that layout, no banned tags (script, foreignObject, iframe, image, use, animation), no on* handlers, no href, no javascript:/data:/vbscript: schemes, no gradients in the watermark, ≤ 20 000 chars, ≤ 200 elements. Anything strippable is stripped; anything that cannot be stripped rejects the whole response.

The same guard runs at dataset build time (a malformed synthetic row never enters training), at benchmark time, in the CLI test harness, and server-side in api.php, which answers 405 for a non-POST request, 400 for a bad prompt, 422 for a blocked prompt or a failed output check, and 502 when Ollama is unreachable.

from logo_model.nsfw import load_blocklist, match_nsfw
from logo_model.svg_guard import sanitize_model_output

terms = load_blocklist("guard/blocklist_tr.txt", "guard/blocklist_en.txt")
match_nsfw("p0rn0 sitesi için logo", terms)   # -> "porn"
sanitize_model_output(raw_model_reply)        # -> {"ok": False, "error": "..."}

Repository layout

configs/        training configs (0.5B/1.5B, custom trainer and `soup` engine)
data/raw/       seed prompts + specs, and the NSFW-filtered copy
data/generated/ chat-format train/val splits (118 / 6)
guard/          system prompt, TR/EN blocklists, TS reference implementations
logo_model/     the installable guard: nsfw.py, svg_guard.py
scripts/
  dataset/      seed_demo -> nsfw_filter -> prepare (+ augment for the loop)
  training/     train_cpu.py, run_train.ps1
  quantization/ merge_adapter.py
  evaluation/   benchmark.py, render_check.py, auto_train_loop.py
  serving/      serve_model.ps1 (Ollama), test_inference.py (e2e)
tests/          guard unit tests (pure stdlib, no torch needed)
webui/          PHP front end + API proxy + PHP port of the guard
docs/           benchmark report, UI screenshot

Getting started

Requires Python 3.10–3.12. Training needs no GPU. Serving additionally needs Ollama and llama.cpp; the web UI needs PHP 8+.

python -m venv .venv
source .venv/bin/activate            # Windows: .venv\Scripts\activate
pip install -r requirements.txt

torch is the only heavy dependency. If you want the CPU-only wheel instead of the default one, install it first: pip install torch --index-url https://download.pytorch.org/whl/cpu.

1 — Data. The splits are committed, so this step is only needed to regenerate them:

python -m scripts.dataset.seed_demo
python -m scripts.dataset.nsfw_filter --in data/raw/seed_records.jsonl --out data/raw/filtered_records.jsonl
python -m scripts.dataset.prepare     --in data/raw/filtered_records.jsonl --out-dir data/generated

2 — Train. About 2 h 45 min on 6 CPU threads for 3 epochs.

python scripts/training/train_cpu.py --config configs/logo-0.5b-cpu.yaml

--resume-adapter models/logo-qwen-coder-0.5b-lora continues from an existing adapter. Threads are set with LOGO_TRAIN_THREADS (default 6). On Windows, scripts/training/run_train.ps1 sets the BLAS thread environment for you. An alternative one-command engine is available via pip install -r requirements-soup.txt then scripts\training\run_train.ps1 -Engine soup.

3 — Merge and quantise.

python scripts/quantization/merge_adapter.py

# point LLAMA_CPP_CONVERT at llama.cpp's converter, or clone it into tools/llama.cpp
python "$LLAMA_CPP_CONVERT" models/logo-qwen-coder-0.5b-merged \
       --outfile models/logo-svg-q8_0.gguf --outtype q8_0

models/ is gitignored — everything in it is regenerable from the base model:

artifact size produced by
models/logo-qwen-coder-0.5b-lora/ 34 MB train_cpu.py
models/logo-qwen-coder-0.5b-merged/ 942 MB merge_adapter.py
models/logo-svg-q8_0.gguf 506 MB llama.cpp converter

4 — Serve and smoke-test.

scripts\serving\serve_model.ps1                       # ollama create logo-svg
python scripts\serving/test_inference.py --prompt "Kahve dükkanı 'Kahve Diyarı' için minimal logo"

test_inference.py runs the input filter, calls Ollama, runs the output guard, and writes data/generated/preview_*.svg — exit code 2/1/3 tells you which stage failed.

5 — Web UI.

webui\serve.ps1                                       # php -S 127.0.0.1:8080 -t webui

Tick Demo modu to exercise the whole pipeline without Ollama installed.

Environment variables used by the auto-train loop: LLAMA_CPP_CONVERT (path to convert_hf_to_gguf.py), OLLAMA_BIN (path to the ollama executable).

Evaluation

python -m scripts.evaluation.benchmark --label manual   # 28 prompts, writes experiments/*.json
python -m scripts.evaluation.auto_train_loop            # full loop, up to 4 cycles
python scripts/evaluation/render_check.py --file data/generated/val.jsonl --render

The loop is benchmark-driven: measure → if any gate fails, augment the dataset → refilter → retrain from the existing adapter → remerge → requantise → redeploy → remeasure. Cycle 1 passed, so the committed model is the original 3-epoch run.

Quality gates are numbers, not vibes: valid_rate ≥ 0.90, brand_rate ≥ 0.85, nsfw_rate = 1.00.

Tests and lint

The guard is pure standard library, so CI runs it without installing torch:

pip install -r requirements-dev.txt
ruff check .
pytest

53 tests cover the accept/reject contract of both guard stages — code fences, missing or extra variant keys, wrong viewBox, banned tags and URI schemes, gradient-in-watermark, oversized markup, leet-speak and letter-spaced evasion, and false-positive checks against ordinary Turkish and English prompts.

Known limitations

  • Composition geometry. scripts/dataset/variants.py places the icon group with a fixed translate/scale per layout. Wider motifs can overflow the viewBox in square/vertical/watermark and crowd the wordmark in horizontal. The guard checks structure, viewBox and size — not bounding boxes. Fixing it means deriving the transform from the motif bbox and retraining.
  • Palette adherence is approximate. The model usually lands in the requested colour family but will occasionally substitute a palette it saw more often in training.
  • Truncation is rejected, not repaired. Generation is capped at 4096 tokens; a run that overruns produces unterminated JSON and the guard throws the whole response away. That is the correct behaviour, but it wastes the compute.
  • The data is synthetic. 124 records built from a fixed vocabulary of 16 brands, 8 palettes and 10 icon motifs. The model overfits to that vocabulary and has no notion of real brand guidelines; more diverse seed data is the obvious next step.

License

MIT for the code in this repository.

Model weights are not committed. The base model Qwen/Qwen2.5-Coder-0.5B-Instruct is downloaded from Hugging Face at runtime and carries its own Qwen licence; any output you generate is subject to it. soup-cli is Apache-2.0.

About

CPU-trained LoRA model for SVG logo generation, with local Ollama inference and deterministic output validation

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages