Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion vllm_omni/engine/duplex/session/emitter.py
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,7 @@ def auto_responds(self) -> bool:
extra = getattr(self._ctx.session.config, "extra_body", None)
if not isinstance(extra, dict):
return False
return extra.get("auto_response") is True or extra.get("full_duplex") is True
return extra.get("auto_response") is True

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] This hunk drops or extra.get("full_duplex") is True from shared `SessionEmitt…

This hunk drops or extra.get("full_duplex") is True from shared SessionEmitter.auto_responds(), so extra_body with only full_duplex: true now yields False with no error on duplex paths that call it. Sibling validators still allow that key through: MiniCPM/PersonaPlex only intersect private-key sets that omit full_duplex, and Aura only type-checks the object — those sessions open. Nemotron alone raises via _require_native_full_duplex when auto_response is True is missing. Alias-only clients therefore get a silent behavior change on non-Nemotron models. Reject leftover full_duplex at session open (or in auto_responds) with an explicit migration error, or document that legacy full_duplex is unrecognized and callers must use auto_response.

Evidence: emitter.py:128 return extra.get("auto_response") is True — full_duplex no longer enables auto-respond. Unchanged by this diff, present in the PR-time tree: minicpmo_4_5/duplex/plugin.py:724-731 private_keys = sorted(PRIVATE_RUNTIME_CONFIG_KEYS.intersection(extra_body)) / raise only on private keys (frozenset at :79-97 omits full_duplex). Unchanged by this diff: personaplex/duplex/serving_adapter.py:151-156 same private-key-only pattern (_PRIVATE_RUNTIME_CONFIG_KEYS at :31-40 omits full_duplex). Unchanged by this diff: aura_omni/duplex/plugin.py:441-445 only rejects non-dict extra_body. Diff-touched: nemotron_voicechat/duplex/serving_adapter.py:120-124 enabled = isinstance(extra_body, dict) and extra_body.get("auto_response") is True then raises unless enabled.

Suggestion: if "full_duplex" in extra:
raise ValueError(
"extra_body.full_duplex is removed; set extra_body.auto_response=true"
)
return extra.get("auto_response") is True


def is_stale_model_output(self, payload: dict[str, object]) -> bool:
"""Whether ``payload`` belongs to a turn the session has already moved past."""
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -117,9 +117,7 @@ def _render_tool_response(output: str) -> str:

def _require_native_full_duplex(config: object) -> None:
extra_body = getattr(config, "extra_body", None)
enabled = isinstance(extra_body, dict) and (
extra_body.get("auto_response") is True or extra_body.get("full_duplex") is True
)
enabled = isinstance(extra_body, dict) and extra_body.get("auto_response") is True
if not enabled:
raise ServingRuntimeConfigError(
"Nemotron VoiceChat currently supports model-native full-duplex streaming only; "
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -160,7 +160,7 @@ def create_session_state(self):
def validate_client_extra_body(self, extra_body):
if not isinstance(extra_body, Mapping):
return
if extra_body.get("auto_response") or extra_body.get("full_duplex"):
if extra_body.get("auto_response"):
raise DuplexRuntimeConfigError(
"Qwen requires server VAD or explicit commits; native auto_response is unsupported"
)
Expand Down
Loading