Skip to content

Repository files navigation

Syntax

Writing-voice API. Style extractor + restyle engine. Opt-in voice assist for editorial workflows.

Takes already-verified copy and restyles it in a chosen editorial voice — without changing any fact, number, score, or proper noun. A human reviews every suggestion before anything is accepted or published. Syntax never writes facts. Syntax never publishes.

Status: v1 built and tested locally. Two endpoints live, three built-in voices, custom voice creation, and a fact-verification layer, all running against a live demo.

Syntax demo — extract a style, restyle verified copy, fact guard flags what changed


Two endpoints (implemented, v1)

POST /api/extract-style Takes pasted writing samples. Returns a structured, named, reusable style profile. No per-style retraining.

POST /api/restyle Takes already-verified copy + a style ID. Returns a restyled suggestion. fact_guard runs on every output.


Architecture

Restyle-only. Facts live upstream in the verified copy; Syntax changes voice, not facts. The restyle engine is model-agnostic behind a stable API surface — the endpoints do not change regardless of what powers them underneath.

Engine roadmap:

  • v1 (current plan): frontier model API. Fastest path to a working, deployable endpoint. In this version the deployed endpoint calls an external model API — it is not running an owned model. Stated plainly so the distinction is never blurred.
  • Target: an owned, self-hosted small style-transfer model, so the intelligence is owned rather than rented. The specific approach is an open decision to be settled by evaluation on real editorial copy — not yet chosen. Candidate directions include instruction-conditioned small models (a general small model directed to restyle via a structured style profile) and authorship-embedding models such as TinyStyler (see References), an MIT-licensed, publicly available 800M-parameter model. Before adopting any specific approach, two things need verifying: output quality on editorial/newsroom copy (the TinyStyler paper evaluates Reddit authorship and formality data, not editorial prose, and its authors note the approach may underperform on rarer stylistic choices not captured by the style embeddings), and whether the chosen model runs acceptably on the intended hardware (TinyStyler's published timing was measured on NVIDIA hardware; Apple Silicon compatibility is unverified).

Serving note: models prototyped/trained on Apple Silicon (MLX) cannot run inside a Linux cloud host (Vercel, Railway, standard GPU providers), because MLX targets Apple Silicon only. Serving an owned model in production therefore means either (a) exporting the weights to a Linux-compatible runtime (e.g. a PyTorch/safetensors export, or a quantized GGUF build) and hosting them on a GPU/CPU host, or (b) using a hosted GPU inference endpoint that serves the model and exposes an HTTPS API the Vercel layer calls. Either way this is an unbuilt, non-trivial step — not assumed complete. Provider choice, cost, and cold-start latency all need evaluation before committing.


Style object schema (proposed)

{
  "id": "string",
  "name": "string",
  "description": "string",
  "exemplars": ["string", "string"],
  "rules": {
    "sentence_length": "string",
    "register": "string",
    "vocabulary": "string",
    "avoid": ["string"]
  }
}

This schema is a design proposal. Its exact fields (exemplars vs. an embedding reference) will depend on which engine approach is chosen above.


fact_guard

Extracts numbers, scores, ordinals, and capitalized proper-noun spans from the source copy. Verifies they survive in the restyled output. Flags anything missing, changed, or newly introduced.

Advisory by design — the human reviewer makes the final call. Syntax augments editorial judgment; it does not replace it.

Known limitation: name detection is a capitalization heuristic, not full named-entity recognition. It will miss lowercase names and may over-flag sentence-initial words. Upgrade path is a real NER model.


Corpus policy

Default built-in styles use public-domain sources only. No named individual's published work is used as training data or exemplar corpus without their explicit written consent.


Prior art

  • PlainSpeak — SmolLM2-1.7B LoRA fine-tuned on dense→plain English, on Apple Silicon.
  • CivicDigest — Small model fine-tuned (LoRA on Apple Silicon) for city-council-minutes summarization.
  • Syntax — style-transfer API, restyle-only, building now.

References

  • Horvitz, Patel, Singh, Callison-Burch, McKeown, Yu. "TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings." Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 13376–13390. https://aclanthology.org/2024.findings-emnlp.781/ — Code (MIT-licensed): github.com/zacharyhorvitz/TinyStyler. Model checkpoints on Hugging Face: tinystyler/tinystyler.

Stack

  • Next.js / Vercel (API surface)
  • Restyle engine: frontier model API (v1) → owned, self-hosted small model (target; approach TBD)
  • Model prototyping/training: Apple Silicon / MLX (M1)
  • MIT License

About

Writing-voice API. Style extractor + restyle engine. Opt-in editorial voice assist.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages