Skip to content

Feature/smooth audio waveforms - #173

Closed
MichaelRicks wants to merge 121 commits into
Lightricks:mainfrom
MichaelRicks:feature/smooth-audio-waveforms
Closed

MichaelRicks wants to merge 121 commits into
Lightricks:mainfrom
MichaelRicks:feature/smooth-audio-waveforms

Conversation

@MichaelRicks

Copy link
Copy Markdown

No description provided.

SunstoneNC and others added 30 commits June 27, 2026 03:01
Integrate the "Prompt Manager Pro" tool into the LTX Desktop fork.

gpm-core (frontend/gpm-core/): pure, dependency-free TS engines ported from
the ltx-gpm.jsx prototype — performance (24 emotion families), shot composer,
camera library, workflow builder, and a transport-injected /api/generate
renderer plus the Project->Act->Scene->Beat->Shot IR types.

UI (frontend/components/gpm/): a collapsible dock (Performance Studio, Shot
Setup, Camera, Workflow) plus folder-based Prompts and Images tabs backed by
IndexedDB, schema-compatible with the Grok extension's v1.6 backup (import/
export, prompt thumbnails). Images can be sent or drag-dropped into the LTX
Gen Space input image; Workflow slots bind library images; Inject pushes the
assembled prompt (and first bound image) to the Gen Space prompt box. Adds a
Clear-prompt control and a left-edge close tab.

Integration seams (no entanglement with LTX view internals): new ProjectContext
channels (genSpacePromptInjection, genSpaceInputImagePath, genSpacePromptClear)
consumed by GenSpace; dock mounted via one line in App.tsx.

Dev isolation: electron/app-paths.ts uses a separate userData folder
(LTXDesktopMikeDev) + app name in dev so the fork runs side-by-side with an
installed LTX Desktop. scripts/launch-fork-dev.cmd backs the desktop shortcut.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The desktop shortcut launched from Explorer couldn't find pnpm (it lives in the
per-user npm global dir, which wasn't reliably on the Explorer-inherited PATH).
Run Vite directly via node from the project's node_modules using absolute paths,
so the launcher depends only on the machine-wide Node install and the project on
disk.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Port the extension's CAM_LIBRARY into gpm-core: 90 cards across 7 categories.
Movement (19) renders pre-rendered inline SVG diagrams; the other categories
(angles, sizes, framing, lighting, film stocks, movie looks — 71) render jpg
reference thumbnails bundled under frontend/assets/gpm-camera/.

- camera.ts: regenerated CAM_CARDS; CameraCard gains optional image/svg; CAM_CATS
  covers all 7 categories.
- CameraPanel: renders a bundled jpg (import.meta.glob), an SVG diagram
  (dangerouslySetInnerHTML for movement), or a gradient fallback. Clicking a card
  injects its phrase as before.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Enlarge/Shrink: every dock tab can expand into a centered modal workspace
  (max-w-3xl, 86vh, blue border) with tab chips to switch while expanded and
  Esc/Shrink to collapse.
- Shot Setup: add a source frame (drag from Images tab, pick from library, or
  upload) with a live 3D-transform preview; WASD+QE drive the virtual camera
  (W/S zoom, A/D rotate, Q/E tilt) alongside the sliders; Reset returns to a
  neutral front view; Apply sends the source image to Gen Space and injects the
  shot prompt.
- Workflow reference-slot thumbnails and the image picker are now 16:9 and
  larger for clearer previews.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Saved Workflows: name + save the current workflow to IndexedDB, with a
  "Saved Workflows (N)" accordion (Load/Delete). Workflows now round-trip in
  the v1.6 backup, mapped to/from the extension's slot shape — so importing an
  existing backup populates the saved list.
- Plates: replace the stub with a panorama library — Import Panorama (file
  picker), drag-to-save (incl. from the Images tab), a 2:1 thumbnail grid with
  delete, capped at 20. Stored in a new IndexedDB 'panoramas' store.
- gpm-storage: DB bumped to v2 (panoramas store); add workflow + panorama
  persistence and backup mapping.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- 3D Scene Builder: port the extension's Three.js panorama viewer into
  public/gpm-scene/ (viewer.html/viewer.js/three.min.js) and embed it via an
  iframe in a new SceneViewer. Open a panorama in 3D, drag to look, pick a lens
  (12-85mm), reset aim, and Generate plate (1280x720) -> sent to Gen Space as
  an input image. Adds a gpm-scene-set-aim message so aim can be restored.
- CSP: frame-ancestors 'none' -> 'self' so the same-origin viewer iframe renders.
- Fix enlarge/shrink dropping panel state: render the active panel once (compact
  body OR modal, not both) and lift ephemeral state to the Dock — Plates active
  panorama + scene lens/aim, and the Shot Setup source image — so enlarging
  preserves the open scene, its lens, and camera aim. Also removes a duplicate
  plate-capture caused by the double-rendered panel.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
A left-side panel backed by real folders/files under <Downloads>/PromptManagerPro,
via new gpmLib* Electron IPC (list/create/rename/delete folder, add/move/delete
file, reveal). Create/rename/delete folders, add images/videos through a native
picker, move media between folders by drag, and open the folder in Explorer.
No FileSystemDirectoryHandle permission dance — real files.

Working: folders, import, move media between folders. Known TODO (next session):
drag-reorder folders, drag a Downloads image into the right-dock Images tab, and
drag a Downloads image onto the Gen Space prompt box; plus save-from-generated
and video thumbnails.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…esent

With an LTX API key, generation can run via the (free) LTX API, so the
required-local-models gate shouldn't force a download to enter the app — it now
resolves to "ready" when a key exists (like force-API mode). Local model
downloads remain available in Settings. This also avoids a cold-boot model-scan
race that wrongly demanded the local gemma text encoder even though the user
uses cloud text encoding.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Performance Studio: dialogue box now mirrors the extension exactly
  (placeholder default per emotion, click-to-promote, one-click clear)
- Downloads Browser: send-to-Gen-Space button, move-to-folder dialog,
  blue drop-line folder reordering, close tab, and main content now
  shifts over instead of covering the editor's Assets panel
- Drag-and-drop: Downloads/Images media can now be dropped onto the
  Gen Space prompt box, the Images tab, the video editor timeline,
  the Clip Viewer, and the Timeline Viewer (auto-copies into project
  storage and probes duration/dimensions)
- Widescreen/video thumbnails, hover preview, and double-click
  lightbox unified across Downloads Browser, Images tab, and workflow
  slots via shared MediaThumb/media-preview components
- Right dock collapsed state is now a side tab instead of a floating
  pill that covered timeline UI; Enlarge button is now a labeled,
  accent-colored primary action
- Tab key collapses/restores both side panels together
- Editor playback now stops at the true end of the last clip (not
  the padded 30s ruler minimum) and respects IN/OUT marks during
  normal playback, not just the dedicated loop button; resuming from
  the OUT point now restarts at IN; added a visible Clear In/Out
  button for the previously shortcut-only action

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…onflict

- Dropping media onto the timeline, Clip Viewer, or Timeline Viewer now
  also registers it in the Assets panel, not just as a timeline clip
- Assets panel now accepts drag-and-drop from the Downloads Browser and
  Images tab directly
- Prompt Manager Pro dock now shifts the editor content over (mirroring
  the Downloads Browser) so it no longer covers the Assets panel
- Raised the GPM panels' z-index above the editor's MenuBar (z-60), which
  was painting over the enlarged Prompt Manager Pro view despite the
  layout shift

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ebrand

- Downloads Browser folders are now true independent accordions (each
  expands inline, multiple can be open) with a thick filled-triangle
  disclosure icon, a Collapse All button, and a type filter (All/Images/
  Videos/Audio)
- Added audio file support end-to-end: real-filesystem library, MediaThumb
  (icon tile + audio lightbox), drag-and-drop payload
- Prompt Manager Pro dock's section chevrons and "collapse all" button now
  match the Downloads Browser's new triangle/icon style
- Gen Space prompt bar's image-ref and audio drop-zone icons doubled in
  size for visibility
- Rebranded the dev fork to "LTX Desktop Studio Pro" (window title, app
  name, launcher script) — userData folder/appId left untouched to avoid
  re-triggering the model-path/sandboxing issues from earlier

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Gen Space (and the Video Editor after first visit) now stay mounted
  across tab switches instead of being torn down — fixes generation
  progress/prompt/result getting silently lost when switching to the
  Video Editor mid-generation
- Downloads Browser: "Set Folder" button to point the media library at
  any folder on disk (persisted across restarts), with a reset-to-default
  action; the chosen folder is added to the app's allowed-roots list
- Gen Space video results: "Post to X" button — opens X's compose page
  with the prompt pre-filled as a caption and reveals the file in
  Explorer so it can be dragged straight into the post (no API/OAuth)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…nlarge modal

- Native asset drag now works directly on the Source/Program monitor drop
  zones, not just double-click.
- Enlarge panel no longer covers the Downloads Browser or dock while open,
  and tab switching no longer closes it on a stray click.
- Drag images from the Images tab into Downloads Browser folders (header
  and open file grid), the Plates panorama list, and Workflow reference
  slots (which also accepts files dragged from Downloads Browser).
- Removed redundant folder icon in Images tab folder headers.
Video Editor's one-time asset snapshot was silently overwriting newer
Gen Space generations on autosave once both views started staying mounted
across tab switches. Merge external assets into the editor model on write
and on live sync, and flush autosave on app close / on demand instead of
relying solely on the debounce timer.

Also adds genuine project portability: Save/Export buttons in the project
header, Import on the Home screen, and matching items in a custom File
menu (New/Save/Export/Import Project) for users who reach for the menu
bar out of habit.
…e Help menu links, sequential timeline names, safer project import

- Rename Downloads panel to "Studio Assets" and right-dock panel to "Desktop Studio Pro"
- Fix Gen Space prompt bar/preview centering to account for open side panels (App.tsx content wrapper now shrinks instead of overflowing)
- Restore Documentation and Keyboard Shortcuts links in the native Help menu
- New timelines get sequential default names (Timeline 2, 3, ...) instead of always "Timeline 1"
- Auto-suffix imported project names on collision to avoid duplicate-looking entries in Recent Projects
…rformance Studio emotions, simplify default tracks

- Route playback audio through a Web Audio GainNode so clip volume can exceed the 100% ceiling HTMLMediaElement.volume imposes; raise UI/clamp limits to 200% (export already supported boosted gain)
- Preview now holds the last clip's final frame when the playhead is past the end of the timeline instead of showing a black "No clip at playhead" screen
- Add 7 emotion families from the original Grok extension (Flirtation, Coyness, Sarcasm, Suspicion, Vulnerability, Jealousy, Embarrassment) with full waypoints and default dialogue; Controls section now re-expands every time the Performance tab is opened
- New timelines default to one video + one audio track instead of three video + two audio, so the first clip lands visibly at the top instead of requiring a scroll
Reuses the existing (previously unused) Asset.binId/project.bins schema fields
to let users file generated images/videos into named folders, similar to Grok
Imagine's tagging UI but with an actual "Untagged" default view so the main
canvas doesn't stay cluttered with everything forever.

- Chip bar (Untagged / All / per-folder / + New Tag) above the asset grid
- Tag dropdown on each asset card to file into a folder, create a new one, or untag
- Deleting a folder un-tags its assets instead of deleting them
- Fixed two bugs found while testing: window.prompt() is a no-op in Electron's
  renderer (replaced with a real modal), and AssetCard's overflow-hidden was
  clipping the tag dropdown invisible (SettingsDropdown now portals its panel
  to document.body instead of relying on position:absolute containment)
…shot

updatedProject() already merged in newly-added assets to avoid clobbering
Gen Space generations with the editor's stale background snapshot, but still
overwrote project.bins wholesale with the editor's copy. Any folder created
in Gen Space's new tagging UI while the editor sat mounted in the background
got silently erased on the editor's next autosave. Union the two instead.
In/out marks lived only in transient session state, never in the saved
Timeline schema, so they silently reset on every reload. Added optional
inPoint/outPoint fields to Timeline, seed the in-memory map from the saved
timeline on mount, and fold the live map back in on every save (autosave and
the explicit commitToProject path) while keeping them out of undo/redo, same
as before — they're playback marks, not document content.
Reported as "nothing happens" on a different machine. handleExportProject
had no try/catch, so any thrown error became a silent unhandled promise
rejection with zero user feedback. Now every failure path shows a toast.
Also logs the one main-process case (missing window reference) that was
previously indistinguishable from a user-cancelled dialog.
Duplicate clones a project with a fresh id/name, referencing existing
asset files rather than copying them, for incremental branching work.

Gen Space thumbnails now use auto-fill grid columns with a fixed min
width instead of viewport-breakpoint column counts, so opening/closing
side panels no longer squeezes thumbnails (and hides the tag icon).
Video Editor's separate thumbnail grid is unaffected.

Co-Authored-By: Claude <noreply@anthropic.com>
The settings.json round-trip for the models dir (app boot -> frontend
sync -> save-on-exit) has proven unreliable in practice. An explicit
env var set by the dev launcher gives a hard guarantee that doesn't
depend on that sync path working.

Co-Authored-By: Claude <noreply@anthropic.com>
Bump auto-fill min thumbnail widths — the previous values (150/220/340px)
were smaller than what the old breakpoint grid rendered at typical
widths, shrinking thumbnails and squeezing the tag icon off the hover
button row. Also let the button row wrap instead of overflow-clipping,
so buttons can't get hidden again regardless of card width.

Co-Authored-By: Claude <noreply@anthropic.com>
Clear Prompt button lets users clear the prompt box without opening
the side panel, matching the existing Prompt Manager Pro action.

Fix Video Editor's background autosave silently resurrecting assets
deleted from Gen Space: updatedProject() merged in the editor's stale
asset snapshot without accounting for assets removed from the live
project since the editor mounted, so a deleted video would reappear
the next time autosave flushed. Now drops assets no longer present on
the live project instead of carrying the stale snapshot forward.

Co-Authored-By: Claude <noreply@anthropic.com>
Bumps the pinned diffusers git rev to 7104cb43c (merge commit for
huggingface/diffusers#14045, "Add Krea 2 (K2) text-to-image pipeline
and transformer"), the earliest commit with Krea2Pipeline /
Krea2Transformer2DModel available. No stable diffusers release
includes it yet (still unreleased as of 0.38.0).

Full backend test suite (257 tests) passes unchanged on this bump.
Manual verification of Z-Image-Turbo and LTX video generation still
to follow before writing any Krea2-specific code.

Co-Authored-By: Claude <noreply@anthropic.com>
Wires Krea 2 Turbo in alongside Z-Image-Turbo as a second selectable
text-to-image checkpoint:

- api_types: add "krea-2-turbo" to ModelCheckpointID, new narrower
  ImageGenerationModelCheckpointID for the two valid image models, and
  a model field on GenerateImageRequest (defaults to z-image-turbo).
- model_download_specs: download spec for krea/Krea-2-Turbo, pointing
  at the already-downloaded Krea-2-Turbo folder.
- New Krea2ImageGenerationPipeline, mirroring ZitImageGenerationPipeline's
  create()/generate()/to() shape against diffusers' Krea2Pipeline.
- pipelines_handler: image_generation_pipeline_class becomes a
  cp_id -> pipeline class dispatch table. load_image_generation_pipeline_to_gpu
  now takes the requested checkpoint id and evicts/reloads on a model
  switch instead of silently reusing whatever was cached.
- image_generation_handler: threads req.model through to pipeline
  loading; Krea 2 is self-hosted only, so force_api_generations with
  krea-2-turbo returns a clear error instead of falling back to the
  fal.ai Z-Image endpoint.

Full backend test suite (257 tests) passes. Frontend model selector
still to come.

Co-Authored-By: Claude <noreply@anthropic.com>
Replaces the static "Z-Image Turbo" indicator in the image-mode prompt
bar with a model dropdown (Z-Image Turbo / Krea 2 Turbo), matching the
existing video model selector pattern. Selection flows through
GenerationSettings.imageModel -> generateImage's request payload's
new `model` field. Krea 2 defaults to 8 inference steps (turbo model
card recommendation) vs Z-Image's 4.

Regenerated OpenAPI schema/types to pick up the backend's new
GenerateImageRequest.model field.

Co-Authored-By: Claude <noreply@anthropic.com>
- Bump transformers to >=5.2 (checkpoint requires 5.x; loads with zero
  workarounds there) and pin diffusers to the rev validated standalone.
- Evict video-family pipelines from the GPU slot before loading an image
  model, and clean up BEFORE creating the new pipeline, not after.
- TorchCleaner: gc.collect() before empty_cache so cyclic pipeline refs
  actually release their VRAM.
- Krea 2 wrapper: group offloading with synchronous copies and
  low_cpu_mem_usage (Windows cannot pin 24 GB), per-step and post-run
  allocator cache flushes, and a ~1 MP generation cap with Lanczos
  upscale to the requested size - the allocator reservation grows toward
  the full transformer size per pass, and larger canvases spill past
  24 GB VRAM into system RAM (~40x slower).

Verified: 256 backend tests + pyright clean; in-app LTX video,
Z-Image-Turbo, and Krea 2 all generate correctly, including
Krea2->video->Krea2 model swaps.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…print

Load the transformer in 4-bit (bitsandbytes NF4) on CUDA instead of the
full bf16 weights streamed via custom block-level offloading. The
quantized transformer (~7 GB) fits on a 24 GB GPU as a whole component,
so diffusers' plain enable_model_cpu_offload replaces the manual
group-offload/pinning workarounds entirely.

Measured on an RTX 3090 at the app's real 1920x1080 request size:
~4.2s/step steady state (was ~11-14s/step), and VRAM reservation drops
to 0 GB when idle instead of staying resident at ~24-35 GB - the
quantized weights should coexist with the video pipeline in system RAM
far more cheaply, shrinking the image<->video swap tax.

The ~1 MP generation cap (upscaled to the requested size) stays: it
addresses activation memory, which scales with pixel count regardless
of weight quantization, not transformer size.

Verified: 256 backend tests + pyright clean; standalone quantization
prototype and the fork's own wrapper both tested directly (two
back-to-back full-resolution generations, image quality unchanged).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Quantization only shrinks the transformer in memory - every cold load
still re-read the full ~24 GB bf16 shards from disk. Under RAM pressure
(e.g. after a large video generation evicts Krea 2's files from the OS
page cache), that turned into a multi-minute stall on the next Krea 2
load even though generation itself was already fast.

Save the quantized transformer to model_path/_nf4_transformer_cache
after the first quantize-from-source load (never touches the original
files) and load from there on subsequent creates. Falls back to
re-quantizing from source if the cache is missing or fails to load.

Verified via the fork's own wrapper: first load 43.0s (quantizes +
builds a 6.8 GB cache), second load 7.5s (5.7x faster), generation
unaffected (~4.2s/step, correct output size and quality). 256 backend
tests + pyright clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
SunstoneNC and others added 28 commits September 16, 2026 21:35
Exporting a timeline with a text overlay failed with "Error parsing filterchain
… No option name near '/Windows/Fonts/…'". The drawtext fontfile was a Windows
drive-letter path whose colon broke ffmpeg's filter parser.

The only escaping that parses in a filter_complex_script on Windows is a
single-quoted path with the colon escaped: fontfile='C\:/Windows/Fonts/arial.ttf'
(bare and backslash-only both fail — verified against the bundled ffmpeg).
Resolve a system SANS font (Arial first, matching the preview's font stack) so
exported overlays look like the editor preview instead of ffmpeg's default serif;
fall back to no fontfile (default font) if none is found. Verified end to end:
the overlay parses and renders in Arial with its shadow.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add audioFadeIn/audioFadeOut (seconds, optional) to clips and apply a linear
gain envelope in both the export mixer and the live preview so they match.

- Export (audio-mix.ts): ramp each source's samples up over audioFadeIn and down
  over audioFadeOut, threaded through ExportClip + the exportNative IPC.
- Preview (usePlaybackAudioSync.ts): multiply the same envelope into each audio
  clip's Web Audio GainNode, at both the playing tick and the paused/scrub path,
  so fades are audible while editing.
- UI: Fade In / Fade Out sliders on audio clips in the ClipPropertiesPanel,
  capped at half the clip length.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Draggable fade-in/fade-out knobs + fade-wedge shading on audio clips
- Audio cross-dissolve at a butted audio seam (hover ⇄ Crossfade button,
  or drop a transition chip), reusing the per-clip audioFade fields so it
  bakes into preview + export for free; drag the ⇄ marker to resize it
- Rubber-band volume line per audio clip: drag up/down for 0–200% gain
  (amber above unity), wired through setClipAudioLevel

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Draw a volume envelope on an audio clip: double-click the flat volume
line to start keyframing, double-click the curve to add points, drag
dots to shape gain (0-200%) over time, double-click a dot to remove.
Piecewise-linear interpolation, composed on top of the fade envelope.

Wired through model (volumeKeyframes), Web Audio preview
(usePlaybackAudioSync), and the PCM export mixer (audio-mix), plus the
export schema/plumbing (ExportModal, electron-api-schema, timeline).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add Fade In / Fade Out sliders to the text overlay properties panel,
defaulting to 0.5s each (a soft fade is almost always wanted; set 0 for
a hard cut). Opacity rides the fade in both the live preview and export.

- Model: textFadeIn/textFadeOut on the clip (optional, capped at half
  the clip); shared textFadeMultiplier helper
- Preview: inline opacity for scrub/paused + a per-frame raf during
  playback (the frame loop short-circuits on unchanged clip sets)
- Export: drawtext alpha='max(0,min(1,...))' ramp, validated against the
  bundled ffmpeg; plumbed via ExportModal -> IPC schema -> video-filter

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Six punch-list fixes lost or broken in the LTX 2.5 migration:

- Image cancel: wire Krea 2 into the interrupt callback. Its generate()
  was the only image path not passing callback_on_step_end, so Stop set
  the interrupt Event but nothing polled it during inference — the run
  pegged the GPU to completion and only an app exit freed it. Z-Image
  already had the hook.
- Final-frame drop: give handleLastDrop the same GPM / Downloads-Browser
  source branches (and object-URL fallback) as the first-frame drop, so
  in-app library drags land on the last frame instead of silently no-op'ing.
- Multi-Angle "Add reference image": snapshot the FileList before clearing
  the input — reading e.target.files after value='' saw zero files.
- Prompt-folder visual cue: restore drag-to-attach on prompt cards
  (setPromptThumb was orphaned; only display + delete survived), with a
  downscaled thumbnail and a remove control.
- Lightbox: add a visible Download button (save image/video) to the
  enlarged asset viewer.
- Image render timer: record renderMs for images (video already did) and
  show it in the enlarged viewer plus a small badge on gallery thumbnails.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Z-Image Turbo loaded raw bf16, whose ~24.6 GB transformer alone fills a
24 GB card. enable_model_cpu_offload()'s whole-module move then overflowed
into shared system RAM and thrashed forever — uninterruptible, since cancel
is only polled between denoise steps. Port the Krea 2 pattern: NF4-quantize
the transformer (ZImageTransformer2DModel) and Qwen3 text encoder via
bitsandbytes with a disk cache beside the model, CUDA-only (MPS/CPU keep
bf16). Add the same ~1 MP activation cap and per-run allocator cleanup.
Verified on the 3090: peak VRAM 24 GB-overflow -> 6.3 GB, ~12s/gen, clean
output; cached reload ~5s.

Krea 2 measured at ~4.17s/step (NF4 compute-bound), so drop its default
from 8 to 6 steps for ~25% faster renders; add named IMAGE_STEPS_KREA2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The grid cards' render-time chips are z-10, and the grid scroll container
establishes no stacking context, so the chips bubbled to the root and
painted over the floating prompt panel (which had no z-index) when it
scrolled under them or the textarea was enlarged. Give the panel z-20 so
its opaque background covers the chips; stays below the z-50 dropdowns and
lightbox so those still layer on top.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
IC-LoRA runs on the full 22B base and can take minutes, but the Output
panel had no way to abort — the only escape was killing the app. Surface a
Stop button under the "Generating…" spinner that calls the existing
process-wide cancel (POST /api/generate/cancel); the IC-LoRA backend
already polls that interrupt between denoise steps. Shows a disabled
"Stopping…" latch after click since cancel lands on the next step boundary.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Three quality-of-life fixes for the Ingredients IC-LoRA workflow:

- Default the IC-LoRA picker to Ingredients (vs Canny Edges) when entering
  IC-LoRA mode from the mode dropdown and it's downloaded — saves a click.
  The panel remounts on every entry and its reset routine told the parent
  to reset conditioning to Canny even when a catalog LoRA was selected,
  deselecting it and wiping its settings on re-entry; guard that reset
  behind !isCatalogIcLora so the selection and settings now persist.
- Tune the Ingredients optimal settings in the catalog default_settings
  (Stage 2 off, LoRA-in-S2 on, RES x2.0, LoRA strength 1.4) so they apply
  on every selection instead of being re-set by hand each run.
- Accept image drag-and-drop into the Ingredients window (input and
  reference) from the Prompt Manager and Studio Assets / Downloads Browser,
  matching the drop sources the rest of Gen Space already supports.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Quick-saving a generation into a Studio Assets folder used to keep its hashed
filename, so themed folders filled with unreadable, unsearchable names. Name
files from the prompt at save time instead: "A toy robot sitting on a shelf..."
saved to Halloween lands as toy-robot-01.png.

- deriveSubject(prompt): a fast, model-free heuristic that takes the opening
  noun phrase (dropping leading framing words and trailing action). Unit-tested.
- gpmLibAddFiles gains an optional baseName: with it, copies land as
  "<subject>-NN.<ext>" sequenced per destination folder (append-only — the
  folder is the sequence source of truth, existing NN are never renumbered);
  without it (imports / drag-drop) the original name is kept.
- The Download quick-save passes the asset prompt; the right-click Save dialog
  now defaults to the clean subject instead of the hash / full prompt.

Steady state needs no vision model — the app already holds the exact prompt at
save time (NSFW included). The synonym-collapsing lexicon is a later phase.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ing fix

- Rename the tool to "Photo Studio + Multi-Angle" (tab + panel header) to
  reflect how much more it does than pose angles.
- Reference thumbnails go from 48x48 to 160x90 (16:9) so they're actually
  legible, with proportionally larger remove / add controls.
- Save now goes through the content-aware auto-naming: derive a subject
  (extra prompt, else the source image's name) and quick-save it as
  subject-NN.png into the last-used Studio Assets folder, same as the Create
  gallery's Download button — instead of qwen_angle_<timestamp>_<full-prompt>.
  Dropped the now-redundant folder picker.
- Fix the extra-prompt box being dead until the first generation: enlarging a
  tab set expandedTab = gpmTab, mounting a second live copy of the same panel
  in the dock that shared state with the enlarged one (double progress polling,
  duplicate inputs fighting over the shared value). The docked strip no longer
  renders a duplicate of whatever tab is enlarged.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Photo Edit + Multi-Angle (Qwen-Image-Edit-2511) improvements:

- Quality modes: replace the binary Lightning toggle with a 3-way
  Fast (4-step) / Balanced (8-step) / Quality (28-step base) selector.
  Adds the lightx2v 8-step Lightning LoRA as a second adapter; skin
  plasticity was the 4-step distillation, so Quality restores detail.

- Role-aware compositing: the References strip is now labeled Location
  and Prop slots. A shared compose_prompt() helper builds an explicit
  instruction ("place the subject from image 1 into the scene shown in
  image 2 ... add the object from image 3 as a prop with the subject")
  so the model composites instead of guessing. Both the live generation
  and the displayed/preview prompt route through the one helper, and the
  client mirror (qwen-angle-mapping.ts composePrompt) keeps the preview
  honest. Prop wording is placement-neutral so it doesn't fight freeform
  text like "slung over her shoulder".

- Rename the tab "Photo Studio + Multi-Angle" -> "Photo Edit + Multi-Angle".

Request gains quality_mode and extra_image_roles (aligned with the refs,
dropped on mismatch). Tests + regenerated OpenAPI types included; pyright,
pytest (TestQwenMultiAngle), and tsc all clean. Live-verified in-app.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Add optional skin-realism adapter (prithivMLmods Hyper-Realistic-
  Portrait, native 2511) as a 4th "skin" LoRA. Its PEFT-style keys are
  normalized (.default. infix dropped, transformer. prefix added) to the
  diffusers convention so every key maps; live-verified applying with no
  missing/unexpected-key warnings. Stacks on any quality mode, weighted
  so it can't overpower the angles LoRA. Exposed as a "Skin realism"
  toggle + Strength slider (0.2-1.5, default 1.0). Restores pore-level
  skin detail on the fast Lightning path instead of needing 28-step.

- Only pass negative_prompt when CFG is on (true_cfg_scale > 1). The
  Lightning modes run at cfg 1.0, where diffusers ignores the negative
  prompt and warns; Quality (cfg 4.0) still gets it.

Request gains use_skin/skin_weight. Tests + regenerated OpenAPI types
included; pyright, pytest (TestQwenMultiAngle), and tsc all clean.
Live-verified in-app (ECU skin detail now production-grade).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Two prompt-generation fixes for the Performance Studio tool:

- Rename the "Subtext:" layer label to "Inner life:". LTX read the
  "Subtext:" line as an on-screen text cue and burned the following
  words into the video. The internal fam.subtext property is unchanged;
  only the emitted label differs.

- Add a Timing line and lead the dialogue-case prompt with a positive,
  present-tense speaking action ("speaking from the first frame"). The
  old negation ("do not hold a long silent pause") fed the model the
  tokens it then rendered, leaving 2-3s of silent performance before
  the first word. Front-loading speech pulls the onset toward frame one.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The Export modal jumped from 0% straight to 100% because ffmpeg progress
was never sent to the renderer (runFfmpeg only logged frame= lines). Wire
real progress end to end:

- runFfmpeg gains an onProgress callback that parses ffmpeg's time= field
  (splitting stderr on \r too, since ffmpeg rewrites its progress line in
  place with carriage returns).
- export-handler computes total program duration and emits an
  export-progress IPC event across stages: Encoding video 0-85%,
  Mixing audio 85%, Finalizing 90-100% (an h264 mux stream-copies and
  snaps to 100%; ProRes/VP9 re-encode and fill the tail for real).
- Expose onExportProgress in preload + schema (ExportProgress type),
  mirroring the onBackendHealthStatus subscription pattern.
- ExportModal subscribes during export, updates the bar and stage label,
  and unsubscribes in finally (guarded by abortRef on cancel).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Inject destination toggle (injectTarget: main | qwen). Reroutes every
  Prompt Manager Inject button (Prompts/Camera/Performance/Shot/Workflow)
  so saved prompts can land in the Multi-Angle extra-prompt box (comma-
  appended, not replaced) instead of the main LTX Create prompt.

- The toggle lives under the extra-prompt textarea inside the panel (easy
  to find where you type), not in the dock headers.

- Auto-switch: opening the enlarged Multi-Angle workspace points injection
  at the panel; shrinking it points back at the LTX prompt box, so text
  can't accidentally be injected into Create. Switching the enlarged view
  to another tab keeps it on qwen so injecting a saved prompt from there
  still lands in the box; manual toggling is respected between transitions.

- Render-time chip (Clock + M:SS, bottom-left) on the Multi-Angle result,
  matching the Create-canvas chips; each generation is timed via
  performance.now() and stored on the result entry.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Prompt visual-cue thumbnails were full-width with object-cover capped at
  120px, so in the enlarged Prompts window they rendered as a long thin
  strip. Bound them to a 16:9 card (max 420px, object-cover) matching the
  Images-tab MediaThumb look; the remove button now anchors to the image.

- Image/prompt folders (Section ids if-*/pf-*) now start COLLAPSED on app
  open instead of expanded, so a large library doesn't fill the panel.
  Functional panel sections (Controls, Presets, ...) still open expanded;
  collapse/expand-all still governs everything (allOverride gates the
  fresh-open default without breaking expand-all).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The enlarged Prompts tab stacked cards in one column while the Images tab
showed two. Thread a `wide` flag through renderPanel (true only for the
enlarged workspace) so PromptsPanel lays cards out in a 2-column grid there,
matching Images; the docked strip stays single-column (too narrow for two).
Also drop the 420px thumbnail cap so visual-cue images fill their grid
column like the Images-tab thumbnails.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the native textarea resize grip (grew downward, felt backwards)
with a custom up-pointing triangle handle above the box: drag up to grow
it, clamped 70px–60vh. Scope a thicker, accent-colored scrollbar to the
prompt box and quiet its hover cursor. Recolor the Clear button red.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a Select mode to the Create-canvas gallery: click cards or drag a
marquee (shift/ctrl to add) to build a selection, then tag or delete
every selected image/video at once. The selection controls sit inline in
the gallery toolbar next to the Select button. Marquee hit-testing runs
in the scroller's content space so it stays correct while scrolling.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Bundles the Gen Space prompt-box resize/polish and library multi-select
(bulk tag + delete) work on top of v1.2.7.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Backend returns the seed it used on image/video complete responses
  (local + FAL image; LTX API video has none).
- Assets persist generationParams.seed (batch image i = seed + i) and
  images get a modelLabel (Krea 2 Turbo / Z-Image Turbo; edits = Z-Image).
- Lightbox footer: model • resolution • duration • Seed N • render time.
- Copy-prompt icon in the top-right of the Gen Space prompt box.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Fix: unlocked seeds were pinned to 1000 in dev mode (_resolve_seed);
  now truly random via secrets.randbelow.
- Prompt bar: Random/Locked toggle, editable seed, re-roll, and
  lock-last-seed button (same app settings as the Settings modal).
- Variation boost slider for local Z-Image Turbo and Krea 2 Turbo:
  seeded noise on the prompt embeddings for the first denoise step,
  clean embeds restored after, so distilled models vary composition
  across seeds while keeping prompt adherence. Deterministic per
  seed+value, no measurable time cost. Shared helper in
  variation_boost.py; Z-Image tuned on a 3090, Krea 2 uses the same
  1-step setting.
- Variation saved to generationParams (lightbox footer shows it);
  Regenerate now restores variation, image model and locks the seed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Krea 2's license requires deployers to run content filters (s4.2) and
forbids circumventing its safety tuning (s4.1c); RiX ships neither, and
Z-Image (Apache 2.0) is faster with comparable results.

- Frontend: Krea 2 option only when import.meta.env.DEV; any Krea
  request (stale setting, Regenerate of a Krea asset) falls back to
  Z-Image in release builds.
- Backend: /api/generate-image rejects krea-2-turbo with
  KREA_2_DEV_ONLY unless dev_mode, so release builds can't run it even
  via a direct API call. Tests for both modes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Natural (masonry) gallery layout: fixed-width columns, cards keep their media's
  true aspect so 9:16 gens show full-size instead of letterboxed in 16:9 cells.
  Ratio from stored dims / image aspect setting, else measured on load (rAF-batched).
  Uniform 16:9 grid kept as a toggle in the grid-size menu (persisted).
- Lightbox insets by open side panels so prev/next arrows are no longer hidden
  under Studio Assets / Prompt Manager; arrows restyled in the accent color.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t newlines

RiX editor MCP server (electron/mcp): lets an MCP client (Claude Code) assemble
videos in the Video Editor from the user's assets.
- Streamable HTTP (stateless JSON) on 127.0.0.1:47821/mcp, bearer token in
  <userData>/rix-mcp.json, Host check, rejects browser Origins. Starts only in
  dev builds or with RIX_MCP=1.
- Tools: rix_status, list_library, list_project_assets, inspect_media (frames +
  soundSegments), analyze_audio (tempo, beat grid, onsets, loudness, sound
  segments), get_timeline, apply_edits, undo, redo, render_frames (fast preview
  render), export_video.
- apply_edits runs against the live editor store (never project files): atomic,
  one undo step. Ops: import, add_clip (in/duration/speed, with_audio,
  audio_track, $ref and $ref.audio), add_text (own title track), update_clip
  (incl. volume_keyframes), duck (music envelope under dialogue ranges),
  transition, delete, clear, new_timeline, add_track.
- Safety: edit/undo/redo require the project_id from rix_status and refuse if
  the user switched projects; undo/redo only walk back the MCP's own batches
  and refuse if anything changed since.
- Main->renderer RPC via mcp-editor-request / mcpEditorResponse; registry kept
  on globalThis so dev hot-reloads don't orphan it.
- Refactors: export payload builder shared by ExportModal and MCP renders
  (export-payload.ts); exportTimelineNative() with a fast preview mode;
  listLibrary() extracted.

Fixes:
- Video Editor imports vanished at the first autosave (since b6106f4): the
  save merge dropped every editor asset absent from the project, unable to
  tell "deleted in Gen Space" from "just imported". Editor-side deletes also
  resurrected. updatedProject now takes the asset ids ever seen on the project
  and ever held by the editor.
- Clip viewer never played audio assets (no media element); audio now plays
  through the hidden <video> so tracks can be auditioned and marked.
- Exported text overlays/subtitles rendered a line break as a literal "n";
  newlines stay raw inside the drawtext value, multi-line text uses text_align
  (textAlign now in the export payload).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Timeline audio clips drew the whole file squeezed into 200 buckets with
nearest-neighbour sampling, so waveforms looked like blocky steps and did
not match the trimmed section that actually plays.

- Decode each file once into a 200 windows/s peak + RMS envelope (all
  channels mixed) and derive every view from it.
- ClipWaveform draws only the clip's source window (trimStart, duration,
  speed, reversed), max-pools when zoomed out and interpolates when zoomed
  in, with a faint peak layer, a bright RMS body and a hairline centre.
- Key the resampled waveform cache by url + bucket count so the monitor
  no longer inherits a coarse resolution from the timeline.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MichaelRicks
MichaelRicks deleted the feature/smooth-audio-waveforms branch September 29, 2026 03:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants