Skip to content
Merged
Show file tree
Hide file tree
Changes from 1 commit
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
65 changes: 65 additions & 0 deletions docs/guides/remove-background.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,71 @@ npx hyperframes remove-background subject.mp4 -o transparent.mov # editi
npx hyperframes remove-background portrait.jpg -o cutout.png # still image
```

## Layer separation: emit the cutout and the background plate together

Pass `--background-output` (alias `-b`) to write a *second* transparent video alongside the cutout. Same source RGB, alpha is the *inverse* mask — opaque where the surroundings were, transparent where the subject is. The result is a clean two-layer separation in a single inference pass:

```bash Terminal
npx hyperframes remove-background subject.mp4 \
-o subject.webm \
--background-output plate.webm
```

| Output | Alpha | Use it as |
| ------ | ----- | --------- |
| `subject.webm` | Mask — subject opaque | Foreground layer (top of stack) |
| `plate.webm` | `255 − mask` — subject region transparent | Background layer; place anything you want **under the subject's silhouette** between this and `subject.webm` |

Both encoders share the source W/H/fps and your `--quality` preset, so the layers are pixel-aligned. Encode cost roughly doubles; segmentation cost is unchanged.

<Tip>
**This is a hole-cut plate, not an inpainted clean plate.** The subject region in `plate.webm` is fully transparent — you have to composite something opaque under it (a graphic, a blurred copy, a different scene) to fill the hole. If you need an actual filled background where the subject was, use a video inpainter (LaMa, ProPainter, RunwayML Inpaint) — `remove-background` is not the right tool for that.
</Tip>

### Hole-cut vs. clean plate — when does the difference matter?

A **hole-cut plate** keeps the original surroundings and makes the subject region transparent. A **clean plate** fills the subject region with reconstructed background — produced by a separate inpainting model. Display each alone over black:

| | Hole-cut plate (this command) | Clean plate (inpainted) |
| --- | --- | --- |
| Subject region | Transparent silhouette | Reconstructed background pixels |
| What you see alone | A person-shaped hole | An empty room |
| Cost | One inference pass, one extra ffmpeg encode | A second model (LaMa, ProPainter, E2FGVI) |
| Tool | `remove-background --background-output` | Outside this CLI |

The line is: **does anything ever need to be visible *through* the subject's silhouette where the subject used to be?**

| Use case | What you need |
| --- | --- |
| Text/graphics live *between* the cutout and the plate (the example above) | **Hole-cut** — the graphics fill the hole. |
| Composite the subject onto an unrelated scene | Neither. Just use `subject.webm`; the plate is irrelevant. |
| Show "the room without the person" as a real background | **Clean plate** — a hole-cut plate would show a transparent void. |
| Replace the person with a different subject (re-target) | **Clean plate** — the new subject needs real pixels under it. |
| VFX rotoscoping / "remove an extra from this take" | **Clean plate** — the canonical inpainting use case. |

If something opaque always covers the silhouette, hole-cut is sufficient and ~1000× cheaper than running an inpainter.

### The two-layer composition pattern

The two-layer pattern is functionally a drop-in for [text-behind-subject](#text-behind-subject-the-recommended-layout) without needing the original `presenter.mp4` in the project — the plate replaces it as the bottom layer:

```html
<!-- z=1 inverse-alpha plate fills everything except the subject's silhouette -->
<video src="plate.webm" data-start="0" data-duration="6" data-track-index="0" muted playsinline></video>

<!-- z=2 anything you want occluded by the subject lives here -->
<h1 style="z-index:2; position:absolute; top:50%; left:50%; transform:translate(-50%,-50%);">
MAKE IT IN HYPERFRAMES
</h1>

<!-- z=3 the cutout puts the subject back on top -->
<div class="cutout-wrap" style="position:absolute;inset:0;z-index:3">
<video src="subject.webm" data-start="0" data-duration="6" data-track-index="1" muted playsinline></video>
</div>
```

Constraints: the flag requires a video input and `.webm` or `.mov` for both outputs. It's not valid for image inputs (no temporal pairing to do) and won't accept `.png` for the plate.

## Performance

Real-world numbers from the [matting eval](https://www.heygenverse.com/a/0dd5a431-1832-4858-862d-de7fb7d02654), running u²-net_human_seg on a 4-second 1080p clip:
Expand Down
7 changes: 6 additions & 1 deletion docs/packages/cli.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -356,6 +356,10 @@ This is suppressed in CI environments, non-TTY shells, and when `HYPERFRAMES_NO_
# Single image → transparent PNG
npx hyperframes remove-background portrait.jpg -o cutout.png

# Layer separation: cutout AND inverse-alpha background plate in one pass
npx hyperframes remove-background avatar.mp4 \
-o subject.webm --background-output plate.webm

# Force CPU on a machine that has CoreML or CUDA
npx hyperframes remove-background avatar.mp4 -o transparent.webm --device cpu

Expand All @@ -366,8 +370,9 @@ This is suppressed in CI environments, non-TTY shells, and when `HYPERFRAMES_NO_
| Flag | Description |
|------|-------------|
| `--output, -o` | Output path. Format inferred from extension: `.webm` (default), `.mov`, `.png` |
| `--background-output, -b` | Optional second output: inverse-alpha background plate (subject region transparent, surroundings opaque). Same source RGB, complementary mask. Must be `.webm` or `.mov`. Hole-cut, not inpainted — composite something underneath to fill the hole. |
| `--device` | Execution provider: `auto` (default), `cpu`, `coreml`, `cuda` |
| `--quality` | WebM encoder preset: `fast` (crf 30, smallest), `balanced` (crf 18, default), `best` (crf 12, near-lossless). Higher quality keeps the cutout's RGB closer to the source mp4 — important when overlaying the cutout on its own source for text-behind-subject effects. Ignored for `.mov` / `.png`. |
| `--quality` | WebM encoder preset: `fast` (crf 30, smallest), `balanced` (crf 18, default), `best` (crf 12, near-lossless). Higher quality keeps the cutout's RGB closer to the source mp4 — important when overlaying the cutout on its own source for text-behind-subject effects. Applies to both `--output` and `--background-output`. Ignored for `.mov` / `.png`. |
| `--info` | Print detected execution providers and exit (no render) |
| `--json` | Output result as JSON |

Expand Down
69 changes: 58 additions & 11 deletions packages/cli/src/background-removal/inference.ts
Original file line number Diff line number Diff line change
Expand Up @@ -24,10 +24,24 @@ interface OrtModule {
Tensor: typeof Tensor;
}

export interface SessionResult {
/** Subject opaque, background fully transparent. */
fg: Buffer;
/** Inverse-alpha plate: same RGB, alpha is `255 − mask`. Null unless `withBackground` was true. */
bg: Buffer | null;
}

export interface Session {
/** Run inference on one RGB frame, return RGBA bytes (H*W*4). */
process(rgb: Buffer, width: number, height: number): Promise<Buffer>;
/** ORT EP that was actually selected. */
/**
* Both `fg` and `bg` (when requested) are session-owned buffers reused on the
* next call — drain the encoder's stdin before invoking `process` again.
*/
process(
rgb: Buffer,
width: number,
height: number,
withBackground?: boolean,
): Promise<SessionResult>;
provider: string;
close(): Promise<void>;
}
Expand Down Expand Up @@ -73,16 +87,15 @@ export async function createSession(options: CreateSessionOptions = {}): Promise
throw new Error("ONNX session is missing input or output bindings");
}

// Pre-allocated per-frame buffers reused across every process() call.
// At 1080p this saves ~9 MB of allocations per frame. rgbaBuf is sized
// lazily on the first call (we don't know W/H until then).
// Reused across calls; sized lazily on first frame. Saves ~9 MB/frame at 1080p.
const inputData = new Float32Array(3 * INPUT_PLANE);
const maskBuf = Buffer.allocUnsafe(INPUT_PLANE);
let rgbaBuf: Buffer | null = null;
let rgbaBgBuf: Buffer | null = null;

return {
provider: providerUsed,
async process(rgb, width, height) {
async process(rgb, width, height, withBackground = false) {
const tensor = await preprocess(sharp, ort, rgb, width, height, inputData);
const outputs = await session.run({ [inputName]: tensor });
const output = outputs[outputName];
Expand All @@ -91,7 +104,21 @@ export async function createSession(options: CreateSessionOptions = {}): Promise
if (!rgbaBuf || rgbaBuf.length !== expectedBytes) {
rgbaBuf = Buffer.allocUnsafe(expectedBytes);
}
return await postprocess(sharp, output, rgb, width, height, maskBuf, rgbaBuf);
if (withBackground) {
if (!rgbaBgBuf || rgbaBgBuf.length !== expectedBytes) {
rgbaBgBuf = Buffer.allocUnsafe(expectedBytes);
}
}
return await postprocess(
sharp,
output,
rgb,
width,
height,
maskBuf,
rgbaBuf,
withBackground ? rgbaBgBuf : null,
);
},
async close() {
await session.release();
Expand Down Expand Up @@ -141,7 +168,8 @@ async function postprocess(
height: number,
maskBuf: Buffer,
rgbaBuf: Buffer,
): Promise<Buffer> {
rgbaBgBuf: Buffer | null,
): Promise<SessionResult> {
const raw = output.data as Float32Array;

let lo = Infinity;
Expand Down Expand Up @@ -172,11 +200,30 @@ async function postprocess(
.raw()
.toBuffer();

for (let i = 0; i < width * height; i++) {
const pixels = width * height;
if (rgbaBgBuf) {
for (let i = 0; i < pixels; i++) {
const r = rgb[i * 3]!;
const g = rgb[i * 3 + 1]!;
const b = rgb[i * 3 + 2]!;
const m = fullMask[i]!;
const o = i * 4;
rgbaBuf[o] = r;
rgbaBuf[o + 1] = g;
rgbaBuf[o + 2] = b;
rgbaBuf[o + 3] = m;
rgbaBgBuf[o] = r;
rgbaBgBuf[o + 1] = g;
rgbaBgBuf[o + 2] = b;
rgbaBgBuf[o + 3] = 255 - m;
}
return { fg: rgbaBuf, bg: rgbaBgBuf };
}
for (let i = 0; i < pixels; i++) {
rgbaBuf[i * 4] = rgb[i * 3]!;
rgbaBuf[i * 4 + 1] = rgb[i * 3 + 1]!;
rgbaBuf[i * 4 + 2] = rgb[i * 3 + 2]!;
rgbaBuf[i * 4 + 3] = fullMask[i]!;
}
return rgbaBuf;
return { fg: rgbaBuf, bg: null };
}
56 changes: 55 additions & 1 deletion packages/cli/src/background-removal/pipeline.test.ts
Original file line number Diff line number Diff line change
@@ -1,7 +1,13 @@
import { describe, expect, it } from "vitest";
import { EventEmitter } from "node:events";
import type { spawn } from "node:child_process";
import { inferOutputFormat, inferInputKind, buildEncoderArgs, waitForExit } from "./pipeline.js";
import {
inferOutputFormat,
inferInputKind,
buildEncoderArgs,
resolveRenderTargets,
waitForExit,
} from "./pipeline.js";

describe("background-removal/pipeline — inferOutputFormat", () => {
it("maps .webm → webm", () => {
Expand Down Expand Up @@ -97,6 +103,54 @@ describe("background-removal/pipeline — buildEncoderArgs", () => {
});
});

describe("background-removal/pipeline — resolveRenderTargets", () => {
it("resolves a normal video → webm render", () => {
const t = resolveRenderTargets("/tmp/clip.mp4", "/tmp/cutout.webm");
expect(t.format).toBe("webm");
expect(t.inputKind).toBe("video");
expect(t.bgFormat).toBeUndefined();
});

it("resolves an image → png render", () => {
const t = resolveRenderTargets("/tmp/portrait.jpg", "/tmp/cutout.png");
expect(t.format).toBe("png");
expect(t.inputKind).toBe("image");
});

it("rejects image input with a video output extension", () => {
expect(() => resolveRenderTargets("/tmp/portrait.jpg", "/tmp/cutout.webm")).toThrow(
/Image input requires a \.png output/,
);
});

it("rejects video input with a .png output", () => {
expect(() => resolveRenderTargets("/tmp/clip.mp4", "/tmp/cutout.png")).toThrow(
/Video input requires a \.webm or \.mov output/,
);
});

it("threads background-output format through when valid", () => {
const t = resolveRenderTargets("/tmp/clip.mp4", "/tmp/fg.webm", "/tmp/bg.webm");
expect(t.bgFormat).toBe("webm");
const tMov = resolveRenderTargets("/tmp/clip.mp4", "/tmp/fg.webm", "/tmp/bg.mov");
expect(tMov.bgFormat).toBe("mov");
});

it("rejects --background-output for image inputs (no temporal pairing to do)", () => {
expect(() =>
resolveRenderTargets("/tmp/portrait.jpg", "/tmp/cutout.png", "/tmp/bg.png"),
).toThrow(/--background-output is not supported for image inputs/);
});

it("rejects .png as the --background-output extension", () => {
// .png is only valid for single-image inputs, and image inputs themselves
// can't have a background-output anyway. So .png here is always a misuse.
expect(() => resolveRenderTargets("/tmp/clip.mp4", "/tmp/fg.webm", "/tmp/bg.png")).toThrow(
/--background-output must be \.webm or \.mov/,
);
});
});

// Regression: a previous version of waitForExit treated `code === null` as
// success. Per Node's child_process docs, that's the signal-killed case —
// reporting it as success means a SIGTERM/SIGKILL'd ffmpeg encoder produces
Expand Down
Loading
Loading