Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
65 commits
Select commit Hold shift + click to select a range
b02905a
workflows comfy
giurgiur99 Aug 4, 2026
08f1e0d
fix template
giurgiur99 Aug 4, 2026
248fe1d
fix download
giurgiur99 Aug 4, 2026
dc783b0
split in two bundles
giurgiur99 Aug 5, 2026
233942c
workflow id
giurgiur99 Aug 5, 2026
613b017
multiscene test
giurgiur99 Aug 5, 2026
11fb2ea
multishoot v2
giurgiur99 Aug 6, 2026
6231b06
fix v2
giurgiur99 Aug 6, 2026
c0f74d5
v3 try
giurgiur99 Aug 6, 2026
3866a6e
voice concat
giurgiur99 Aug 6, 2026
989a93a
update ugc product template to include new fields necessary on dashboard
dnsi0 Aug 6, 2026
6e87150
concat voices too
giurgiur99 Aug 6, 2026
434f221
install missing services
giurgiur99 Aug 7, 2026
3fe722e
cut to new scene
giurgiur99 Aug 7, 2026
26f8dde
remove sizeGb
dnsi0 Aug 11, 2026
803503a
cleanup
giurgiur99 Aug 11, 2026
01a6c2d
Merge branch 'feat/ltx-video-ugc-template' of https://github.com/ocea…
giurgiur99 Aug 11, 2026
373bff2
readd sizegb and schema
giurgiur99 Aug 11, 2026
07ec8c7
minimax flow
giurgiur99 Aug 12, 2026
db1dc33
simplify
giurgiur99 Aug 12, 2026
4a18194
minimax h3 v2
giurgiur99 Aug 12, 2026
be813f1
v3 h3
giurgiur99 Aug 12, 2026
f8f3e0b
h3 v4
giurgiur99 Aug 12, 2026
c791959
image defaults
giurgiur99 Aug 12, 2026
c34cb01
speed improvement
giurgiur99 Aug 12, 2026
5486dc9
fixes
giurgiur99 Aug 13, 2026
fcd927f
fix box sizes
giurgiur99 Aug 13, 2026
2c6ac40
fix corrupt model download
giurgiur99 Aug 13, 2026
e83c4bb
new carachters
giurgiur99 Aug 13, 2026
8ff1fe6
fix tail sound
giurgiur99 Aug 13, 2026
93640b7
add prompt creation model
giurgiur99 Aug 14, 2026
8495e25
add minimax-music3 template
dnsi0 Aug 14, 2026
e88e4a4
ltx 2.5
giurgiur99 Aug 17, 2026
9e753f5
Merge branch 'feat/ltx-video-ugc-template' of https://github.com/ocea…
giurgiur99 Aug 17, 2026
e94fd88
space out boxes ltx 2.5
giurgiur99 Aug 17, 2026
adaa5fc
fix template id
giurgiur99 Aug 17, 2026
f992df0
simplify flow
giurgiur99 Aug 17, 2026
a5864f4
fix bugs
giurgiur99 Aug 17, 2026
d847cec
fix wrong img size
giurgiur99 Aug 17, 2026
a5deb5a
fixes from official docs
giurgiur99 Aug 17, 2026
67a8261
fix input out of range
giurgiur99 Aug 17, 2026
e31c579
minimax h3 fix models
giurgiur99 Aug 18, 2026
3234865
fix new carachters
giurgiur99 Aug 18, 2026
d4f7d21
minimax fixes
giurgiur99 Aug 18, 2026
0ae6611
find old videos
giurgiur99 Aug 18, 2026
d00d5c6
fixes ltx 2.5
giurgiur99 Aug 18, 2026
388b3fd
try open github minimax flow
giurgiur99 Aug 18, 2026
60f472a
live preview
giurgiur99 Aug 18, 2026
61fa51d
allinone fixes
giurgiur99 Aug 18, 2026
02c8cfe
portarait mode
giurgiur99 Aug 18, 2026
c02a081
more allinone fixes
giurgiur99 Aug 19, 2026
7fa8d44
fix extend video
giurgiur99 Aug 19, 2026
9edadff
fix libraries
giurgiur99 Aug 19, 2026
9694c95
add muse, qwen38 templates
dnsi0 Aug 19, 2026
b4015ff
fix openweb ui bootstrap
dnsi0 Aug 19, 2026
c10f9c3
remove pre configured admin
dnsi0 Aug 19, 2026
fc47be5
fix muse tag
dnsi0 Aug 19, 2026
9661d15
Merge branch 'next-4' into feat/ltx-video-ugc-template
alexcos20 Aug 20, 2026
f7d347e
ui flows
giurgiur99 Aug 20, 2026
975e9cd
ecommerce flow fixes
giurgiur99 Aug 20, 2026
2ef625e
update templates to run with opencode
dnsi0 Aug 20, 2026
beacb3d
text changes
giurgiur99 Aug 20, 2026
d4a8009
Merge branch 'feat/ltx-video-ugc-template' of https://github.com/ocea…
giurgiur99 Aug 20, 2026
46bae73
fixes ecom
giurgiur99 Aug 20, 2026
e0aba7d
fixes for ecom
giurgiur99 Aug 20, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
712 changes: 711 additions & 1 deletion docs/serviceTemplates/README.md

Large diffs are not rendered by default.

709 changes: 709 additions & 0 deletions docs/serviceTemplates/comfyui-ugc-bootstrap.sh

Large diffs are not rendered by default.

168 changes: 168 additions & 0 deletions docs/serviceTemplates/ecommerce-studio.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,168 @@
{
"id": "ecommerce-studio",
"name": "ComfyUI \u2014 E-commerce studio (product images, angles, creatives, motion)",
"description": "ComfyUI preloaded with Qwen-Image-Edit 2509 and LTX-2.3 behind five workflows that turn one product photo into the asset set a listing needs. Background puts a packshot into a generated scene. Relight matches the light to a reference photo. Angles makes a full 360 of the product at 45-degree steps, plus a close-up and a loop that starts on your own photo. Creative turns the product photo into an ad poster with your headline written on it. Motion makes a 9:16 clip of up to 5 seconds. Pick a workflow in the tab bar, upload one photo, press Queue. The edit backbone is 2509 and not the newer 2511, because the four product LoRAs this template uses exist only for 2509. Weights are 74.87 GB. The 9.4 GB Qwen2.5-VL text encoder is shared by the four image workflows, and one 20.4 GB edit backbone is shared by all four. Motion uses LTX-2.3, the same 43 GB weight set as the ltx-video-ugc-product template: if you launch with the bucket you already used for that template, those files are on disk and do not download again. Only one diffusion model is in VRAM per generation, because ComfyUI unloads between passes. Select a persistent-storage bucket on launch, or the weights download every time and the renders are lost on stop. Outputs save flat to the bucket root with ecom_ prefixes. Needs a CUDA GPU: the image workflows need about 30 GB, and Motion needs about 40 GB, so 48 GB runs everything. A model invents every image these workflows make. Check labels, small print and prices before you publish. At 135, 180 and 225 degrees the Angles workflow invents the side of the product it cannot see.",
"kind": "bundle",
"service": "comfyui",
"outcome": "Turn one product photo into scene shots, relights, a 360, an ad creative and a clip.",
"category": "image",
"includes": [
{
"name": "Qwen2.5-VL 7B text encoder (fp8 scaled)",
"kind": "model",
"sizeGb": 9.38,
"repoId": "Comfy-Org/Qwen-Image_ComfyUI"
},
{
"name": "Qwen-Image VAE",
"kind": "model",
"sizeGb": 0.25,
"repoId": "Comfy-Org/Qwen-Image_ComfyUI"
},
{
"name": "Qwen-Image-Edit 2509 (fp8 e4m3fn)",
"kind": "model",
"sizeGb": 20.43,
"repoId": "Comfy-Org/Qwen-Image-Edit_ComfyUI"
},
{
"name": "Qwen-Image-Edit 2509 White-to-Scene LoRA",
"kind": "model",
"sizeGb": 0.24,
"repoId": "Comfy-Org/Qwen-Image-Edit_ComfyUI"
},
{
"name": "Qwen-Image-Edit 2509 Relight LoRA",
"kind": "model",
"sizeGb": 0.24,
"repoId": "Comfy-Org/Qwen-Image-Edit_ComfyUI"
},
{
"name": "Qwen-Image-Edit 2509 Light-Migration LoRA",
"kind": "model",
"sizeGb": 0.24,
"repoId": "Comfy-Org/Qwen-Image-Edit_ComfyUI"
},
{
"name": "Qwen-Image-Edit 2509 Multiple-Angles LoRA",
"kind": "model",
"sizeGb": 0.24,
"repoId": "Comfy-Org/Qwen-Image-Edit_ComfyUI"
},
{
"name": "Qwen-Image-Edit 2509 Lightning 4-step LoRA (bf16)",
"kind": "model",
"sizeGb": 0.85,
"repoId": "lightx2v/Qwen-Image-Lightning"
},
{
"name": "LTX-2.3 22B dev (fp8)",
"kind": "model",
"sizeGb": 29.2,
"repoId": "Lightricks/LTX-2.3-fp8"
},
{
"name": "Gemma-3-12B-it text encoder (fp4 mixed)",
"kind": "model",
"sizeGb": 9.5,
"repoId": "Comfy-Org/ltx-2"
},
{
"name": "LTX-2.3 22B distilled LoRA (rank 111)",
"kind": "model",
"sizeGb": 2.7,
"repoId": "Comfy-Org/ltx-2.3"
},
{
"name": "LTX-2.3 spatial upscaler x2",
"kind": "model",
"sizeGb": 1.0,
"repoId": "Lightricks/LTX-2.3"
},
{
"name": "Gemma-3-12B-it abliterated LoRA (rank 64)",
"kind": "model",
"sizeGb": 0.6,
"repoId": "Comfy-Org/ltx-2"
}
],
"image": "yanwk/comfyui-boot",
"tag": "cu130-megapak-pt211-20260814",
"exposedPorts": [
8188
],
"entrypoint": [
"/bin/bash",
"-c"
],
"commandFile": "comfyui-ugc-bootstrap.sh",
"workflows": [
{
"id": "ocean_ecom_background",
"name": "Background \u2014 packshot into a scene",
"description": "Upload a product photo on a plain background and describe the scene you want. The model keeps the product's shape, label and detail and rebuilds everything around it. Saves as ecom_background_00001_.png.",
"file": "workflows/ocean_ecom_background.json"
},
{
"id": "ocean_ecom_relight",
"name": "Relight \u2014 match the light to a reference",
"description": "Upload your product plus a photo whose lighting you want, and the product is re-lit to match. Swap one widget to switch to prompt-driven relighting instead. Saves as ecom_relight_00001_.png.",
"file": "workflows/ocean_ecom_relight.json"
},
{
"id": "ocean_ecom_multiview",
"name": "Angles \u2014 a full 360 of the product",
"description": "Upload one photo. Get 7 rotation steps at 45 degrees, a close-up, and a 360 loop that starts on your own photo. Saves as ecom_angle_045 to ecom_angle_315, ecom_closeup, and ecom_turntable_00001_.webp. Check the back views: the model invents the side it cannot see.",
"file": "workflows/ocean_ecom_multiview.json"
},
{
"id": "ocean_ecom_creative",
"name": "Creative \u2014 ad poster with your headline",
"description": "Upload your product photo and describe the background and the headline you want. One pass writes the words and the background around your product. Saves as ecom_creative_00001_.png.",
"file": "workflows/ocean_ecom_creative.json"
},
{
"id": "ocean_ecom_motion",
"name": "Motion \u2014 a 9:16 product clip",
"description": "Upload one photo and describe the motion, not the product. LTX-2.3 makes a 9:16 clip of up to 5 seconds at 25 fps. Saves as ecom_motion_00001_.mp4.",
"file": "workflows/ocean_ecom_motion.json"
}
],
"userConfigurableEnvVars": [
{
"key": "COMFY_WORKFLOW_ID",
"validation": "^[A-Za-z0-9_-]+$"
},
{
"key": "COMFY_WORKFLOW"
}
],
"requiredResources": [
{
"id": "cpu",
"min": 8,
"recommended": 16,
"unit": "cores"
},
{
"id": "ram",
"min": 64,
"recommended": 128,
"unit": "GB"
},
{
"id": "disk",
"min": 110,
"recommended": 180,
"unit": "GB"
},
{
"kind": "discrete",
"type": "gpu",
"min": 1,
"recommended": 1,
"unit": "count",
"description": "One CUDA GPU. A generation holds one 20.4 GB diffusion backbone plus the 9.4 GB text encoder and a 0.25 GB VAE \u2014 about 30 GB \u2014 so a 32 GB card runs every workflow comfortably and 24 GB works with ComfyUI paging the text encoder out between passes. The two image backbones never co-reside: the Creative workflow runs its two stages as separate passes and ComfyUI unloads between them, so the 58.5 GB on disk is not a VRAM figure. The Motion workflow is the lightest, at 4.6 GB of weights plus the shared encoder. Ask for one card and not two: core ComfyUI has no tensor parallelism, a single sampling pass runs on one device, and a second GPU would be billed while idle."
}
]
}
47 changes: 47 additions & 0 deletions docs/serviceTemplates/ltx-2.5-video-ugc-multishot.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
{
"id": "ltx-2-5-video-ugc-multishot",
"name": "ComfyUI — UGC multishot reel (LTX-2.5)",
"description": "ComfyUI preloaded with LTX-2.5 for a vertical UGC reel where one Run is the whole reel. LTX-2.5 cuts between shots inside a single generation and holds character identity, wardrobe, room, lighting and voice across those cuts, so the shot-at-a-time chaining the LTX-2.3 template needs — render a clip, save its last frame, feed it to the next Run, then stitch and voice-convert the result — is gone, and with it the voice-conversion node pack whose first-launch pip install made that template slow to start. Multishot is a prompt format rather than a node: you write the reel as one chronological prose paragraph and mark each cut in words ('A hard cut transitions to a medium close-up of her face'), restating the character's look and saying what the audio does at every cut ('the guitar track continues across the cut'). One prompt box holds that single paragraph: open with the cast and the look, then the shots, so a re-take means editing the shots and leaving the opening lines untouched. LTX's own guidance is 2-4 shots per generation, and the shipped 8 seconds at 24 fps (193 frames) gives each of them 2 to 4 seconds. The prompt enhancer that ships in the upstream graph is deliberately deleted rather than switched off: it rewrites your prompt into a fuller description and flattens exactly the cuts and re-identifications that make a paragraph multishot, and because ComfyUI validates combo widgets across the whole reachable graph — even branches that never execute — leaving it bypassed would still force its 10.28 GB encoder to download. The same logic sets the download list at exactly five files and nothing more: no duration head (it has no ComfyUI node at all, only a ltx-pipelines CLI flag, and it would fight fixed settings anyway), no temporal upscaler, no second text encoder. Settings track Lightricks' own operating point, because that is what quality depends on: their ltx-pipelines DistilledPipeline defaults to 121 frames (5 s at 24 fps) with an 8-step then 3-step sigma schedule, and their ComfyUI API template ships 8 s. This template ships 8 s / 193 frames with stage 1 at 576x1024 doubling to 1152x2048, about 2.3x the official token count — room for a 9:16 frame while still clearing 1080x1920 delivery. It previously ran 481 frames at 768x1344 into 1536x2688, ten times the official token count on the same fixed step budget, which is why it looked softer than the LTX-2.3 template: raising duration and resolution spreads the same 8 plus 3 steps over more tokens instead of adding any. Render two Runs and join them for a longer reel. Attention memory is linear rather than quadratic under the image's SDPA kernels, which is why this fits at all; the genuine memory risk is the VAE decode, handled by the tiled decoder at 768/192/512/0 — spatial tiling only. Splitting the decode temporally returns each overlap window as red/blue dither rather than a blend, so temporal_size stays above the frame count and temporal_overlap stays at 0. Drop duration or stage-1 resolution — both are dials on the canvas — for a cheaper Run. Weights are bf16 throughout and total about 71 GB, and every LTX-2.5 file is behind a licence gate: open each model page once, click 'Agree and Access' with the account that owns your token, and set HF_TOKEN, or nothing downloads and ComfyUI starts empty. Because the weights are large relative to any single card, the launch script adds --highvram only when VRAM covers them with 3/2 headroom, which happens on an H200 and does not on an 80 GB H100; on the smaller card ComfyUI offloads between stages instead, which works but is slower. Ask for one GPU and not two: the two stages are sequential by construction and the distilled model runs CFG=1, so the CFG-split node has a single conditioning and nothing to hand a second device, and core ComfyUI has no tensor parallelism. Select a persistent-storage bucket on launch — it holds ComfyUI's whole base directory, so the 71 GB downloads once instead of on every launch, and finished reels land in the bucket root where the storage API's listFiles can see them.",
"kind": "bundle",
"service": "comfyui",
"outcome": "Render a complete multi-shot 9:16 UGC reel — one character, one voice, native synced audio — in a single Run.",
"category": "video",
"includes": [
{ "name": "LTX-2.5 22B distilled transformer (bf16)", "kind": "model", "sizeGb": 42.0, "repoId": "Lightricks/LTX-2.5" },
{ "name": "Gemma-4-12B text encoder with projections (bf16)", "kind": "model", "sizeGb": 26.3, "repoId": "Lightricks/LTX-2.5" },
{ "name": "LTX-2.5 video VAE / DiffVAE (bf16)", "kind": "model", "sizeGb": 1.5, "repoId": "Lightricks/LTX-2.5" },
{ "name": "LTX-2.5 audio VAE + vocoder (bf16)", "kind": "model", "sizeGb": 0.4, "repoId": "Lightricks/LTX-2.5" },
{ "name": "LTX-2.3 spatial upscaler x2", "kind": "model", "sizeGb": 1.0, "repoId": "Lightricks/LTX-2.3" }
],
"image": "yanwk/comfyui-boot",
"tag": "cu130-megapak-pt211-20260812",
"exposedPorts": [8188],
"entrypoint": ["/bin/bash", "-c"],
"commandFile": "comfyui-ugc-bootstrap.sh",
"workflows": [
{
"id": "ocean_ltx25_ugc_multishot",
"name": "Multishot reel — one Run, 8 s",
"description": "One Run renders the whole reel: 2-4 shots, one character, one voice, native synced audio, 1152x2048 at 24 fps. Write the reel in one prompt box as a single chronological paragraph: open with the cast, wardrobe, room, lighting and camera, then the shots, marking every cut in prose and restating the character's look and the audio's continuity at each one. Optionally upload a character photo — img_strength on Preprocess decides whether it is literally frame 0 (0.7+) or only a look reference (~0.35).",
"file": "workflows/ocean_ltx25_ugc_multishot.json"
}
],
"userConfigurableEnvVars": [
{ "key": "COMFY_WORKFLOW_ID", "validation": "^[A-Za-z0-9_-]+$" },
{ "key": "COMFY_WORKFLOW" },
{ "key": "HF_TOKEN", "validation": "^hf_[A-Za-z0-9]{20,}$", "sensitive": true }
],
"requiredResources": [
{ "id": "cpu", "min": 8, "recommended": 16, "unit": "cores" },
{ "id": "ram", "min": 128, "recommended": 256, "unit": "GB" },
{ "id": "disk", "min": 110, "recommended": 180, "unit": "GB" },
{
"kind": "discrete",
"type": "gpu",
"min": 1,
"recommended": 1,
"unit": "count",
"description": "One CUDA GPU, H200-class (141 GB). The bf16 weight set totals about 71 GB, and stage 2 samples 57,600 latent tokens at 1152x2048 by 193 frames, so the weights rather than the activations are what fills a card here — an 80 GB H100 fits them but without the 3/2 headroom the launch script wants for --highvram, so it offloads between stages and completes slowly rather than failing. Don't set CLI_ARGS yourself: the launch script measures the card's VRAM against the on-disk size of what it downloaded and adds --highvram only when VRAM covers the weights with 3/2 headroom (about 107 GB here), so it takes the flag on an H200 and correctly declines it on an H100. Ask for one card and not two. The two stages are sequential by construction — stage 2 resamples stage 1's latent — and the distilled model is CFG-distilled, so the CFG-split node sees a single conditioning and has nothing to hand a second device; core ComfyUI has no tensor parallelism either. A second GPU helps only by running a second container for a separate job. GPU class matters directly here; CPU cores do not, since sampling is GPU-bound."
}
]
}
Loading
Loading