Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)pip install evaluate rouge_score datasets\n\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **Installed evaluation dependencies:**\n\n```sh\npip install evaluate rouge_score datasets\n```\n\n4. **At least one GPU with CUDA 13+**",
+ "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)pip install evaluate rouge_score datasets\n\nIn this tutorial we use two Llama 3.2 Instruct models from HuggingFace:
\nBoth models share the same tokenizer/vocabulary (required for knowledge distillation) and include a chat template for deployment with /chat/completions.
HuggingFace Authentication:
\ntoken_secret parameterIn this tutorial we use two Llama 3.2 Instruct models from Hugging Face:
\nBoth models share the same tokenizer/vocabulary (required for knowledge distillation) and include a chat template for deployment with /chat/completions.
Hugging Face Authentication:
\ntoken_secret parameterInterpreting ROUGE Scores:
\n| Metric | \nMeasures | \n
|---|---|
| ROUGE-1 | \nUnigram overlap between prediction and reference | \n
| ROUGE-2 | \nBigram overlap (captures phrase-level similarity) | \n
| ROUGE-L | \nLongest common subsequence (captures sentence structure) | \n
| ROUGE-Lsum | \nROUGE-L computed over full summaries | \n
What to expect:
\ndistillation_temperature, adjusting distillation_ratio, or training for more epochsFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\nmodel (student) and teacher_model URNs are correctclient.models.retrieve(name=..., workspace="default")Job fails with OOM (Out of Memory) error:
\nKD loads both models, so OOM is more likely than with SFT:
\nteacher_precision="bf16" to reduce teacher memorymicro_batch_size to 1global_batch_size and max_seq_lengthnum_gpus_per_nodeNo chat template / /chat/completions fails:
Llama-3.2-1B-Instruct) instead of base models (Llama-3.2-1B). Base models do not include a chat template in their tokenizer, so the output model will also lack one.Distilled model quality is poor:
\ndistillation_temperature (try 2.0–5.0) to transfer more nuanced knowledgedistillation_ratio—if dataset labels are high-quality, lower the ratio; if the teacher is strong, raise itepochs or max_steps for more trainingVocabulary mismatch error:
\nDeployment fails:
\nclient.models.retrieve(name=DISTILLED_STUDENT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")Interpreting ROUGE Scores:
\n| Metric | \nMeasures | \n
|---|---|
| ROUGE-1 | \nUnigram overlap between prediction and reference | \n
| ROUGE-2 | \nBigram overlap (captures phrase-level similarity) | \n
| ROUGE-L | \nLongest common subsequence (captures sentence structure) | \n
| ROUGE-Lsum | \nROUGE-L computed over full summaries | \n
What to expect:
\ndistillation_temperature, adjusting distillation_ratio, or training for more epochsFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\nmodel (student) and teacher_model URNs are correctclient.models.retrieve(name=..., workspace="default")Job fails with OOM (Out of Memory) error:
\nKD loads both models, so OOM is more likely than with SFT:
\nteacher_precision="bf16" to reduce teacher memorymicro_batch_size to 1global_batch_size and max_seq_lengthnum_gpus_per_nodeNo chat template / /chat/completions fails:
Llama-3.2-1B-Instruct) instead of base models (Llama-3.2-1B). Base models do not include a chat template in their tokenizer, so the output model will also lack one.Distilled model quality is poor:
\ndistillation_temperature (try 2.0–5.0) to transfer more nuanced knowledgedistillation_ratio—if dataset labels are high-quality, lower the ratio; if the teacher is strong, raise itepochs or max_steps for more trainingVocabulary mismatch error:
\nDeployment fails:
\nclient.models.retrieve(name=DISTILLED_STUDENT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)pip install evaluate rouge_score datasets\n\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **Installed evaluation dependencies:**\n\n```sh\npip install evaluate rouge_score datasets\n```\n\n4. **At least one GPU with CUDA 13+**",
+ "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)pip install evaluate rouge_score datasets\n\nIn this tutorial we use two Llama 3.2 Instruct models from HuggingFace:
\nBoth models share the same tokenizer/vocabulary (required for knowledge distillation) and include a chat template for deployment with /chat/completions.
HuggingFace Authentication:
\ntoken_secret parameterIn this tutorial we use two Llama 3.2 Instruct models from Hugging Face:
\nBoth models share the same tokenizer/vocabulary (required for knowledge distillation) and include a chat template for deployment with /chat/completions.
Hugging Face Authentication:
\ntoken_secret parameterInterpreting ROUGE Scores:
\n| Metric | \nMeasures | \n
|---|---|
| ROUGE-1 | \nUnigram overlap between prediction and reference | \n
| ROUGE-2 | \nBigram overlap (captures phrase-level similarity) | \n
| ROUGE-L | \nLongest common subsequence (captures sentence structure) | \n
| ROUGE-Lsum | \nROUGE-L computed over full summaries | \n
What to expect:
\ndistillation_temperature, adjusting distillation_ratio, or training for more epochsFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\nmodel (student) and teacher_model URNs are correctclient.models.retrieve(name=..., workspace="default")Job fails with OOM (Out of Memory) error:
\nKD loads both models, so OOM is more likely than with SFT:
\nteacher_precision="bf16" to reduce teacher memorymicro_batch_size to 1global_batch_size and max_seq_lengthnum_gpus_per_nodeNo chat template / /chat/completions fails:
Llama-3.2-1B-Instruct) instead of base models (Llama-3.2-1B). Base models do not include a chat template in their tokenizer, so the output model will also lack one.Distilled model quality is poor:
\ndistillation_temperature (try 2.0–5.0) to transfer more nuanced knowledgedistillation_ratio—if dataset labels are high-quality, lower the ratio; if the teacher is strong, raise itepochs or max_steps for more trainingVocabulary mismatch error:
\nDeployment fails:
\nclient.models.retrieve(name=DISTILLED_STUDENT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")Interpreting ROUGE Scores:
\n| Metric | \nMeasures | \n
|---|---|
| ROUGE-1 | \nUnigram overlap between prediction and reference | \n
| ROUGE-2 | \nBigram overlap (captures phrase-level similarity) | \n
| ROUGE-L | \nLongest common subsequence (captures sentence structure) | \n
| ROUGE-Lsum | \nROUGE-L computed over full summaries | \n
What to expect:
\ndistillation_temperature, adjusting distillation_ratio, or training for more epochsFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\nmodel (student) and teacher_model URNs are correctclient.models.retrieve(name=..., workspace="default")Job fails with OOM (Out of Memory) error:
\nKD loads both models, so OOM is more likely than with SFT:
\nteacher_precision="bf16" to reduce teacher memorymicro_batch_size to 1global_batch_size and max_seq_lengthnum_gpus_per_nodeNo chat template / /chat/completions fails:
Llama-3.2-1B-Instruct) instead of base models (Llama-3.2-1B). Base models do not include a chat template in their tokenizer, so the output model will also lack one.Distilled model quality is poor:
\ndistillation_temperature (try 2.0–5.0) to transfer more nuanced knowledgedistillation_ratio—if dataset labels are high-quality, lower the ratio; if the teacher is strong, raise itepochs or max_steps for more trainingVocabulary mismatch error:
\nDeployment fails:
\nclient.models.retrieve(name=DISTILLED_STUDENT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")Learn how to use the NeMo Platform to align a model with DPO (Direct Preference Optimization) on a preference dataset. For each prompt, DPO trains on a chosen (preferred) and a rejected response so the model prefers the chosen style — no separate reward model required.
\nThis tutorial uses the rl customization backend (powered by NVIDIA NeMo-RL), which runs DPO on a Ray cluster. Unlike the SFT and LoRA tutorials (Docker GPU jobs), rl requires a Kubernetes-backed NeMo Platform. DPO here is full-weight (no LoRA/adapter); the output is a full model entity.
Time to complete: approximately 45-60 minutes. Job duration increases with model and dataset size.
\n" + }, + { + "type": "markdown", + "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](/documentation/get-started)** to install the NeMo Platform and Python SDK.\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root).\n3. **Installed the `datasets` package**: `pip install datasets`.\n4. **A platform configured with `platform.runtime: kubernetes`.** The `rl` (DPO) backend provisions a Ray cluster and has **no local Docker fallback** — `submit` fails fast on a Docker-runtime platform. Multi-node jobs (`parallelism.num_nodes > 1`) additionally require the platform-side `NMP_RL_MULTINODE_SHARED_STORAGE_PATH`.\n5. **A Hugging Face token** with access to the gated base model (this tutorial uses `meta-llama/Llama-3.2-1B-Instruct`). Export it as `HF_TOKEN`.\n6. **At least one GPU with CUDA 13+** and a GPU execution profile (`nemo jobs list-execution-profiles`).", + "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root).datasets package: pip install datasets.platform.runtime: kubernetes. The rl (DPO) backend provisions a Ray cluster and has no local Docker fallback — submit fails fast on a Docker-runtime platform. Multi-node jobs (parallelism.num_nodes > 1) additionally require the platform-side NMP_RL_MULTINODE_SHARED_STORAGE_PATH.meta-llama/Llama-3.2-1B-Instruct). Export it as HF_TOKEN.nemo jobs list-execution-profiles).The SDK needs your NeMo Platform server URL. By default http://localhost:8080 is used; set NMP_BASE_URL to override:
export NMP_BASE_URL=<YOUR_NMP_BASE_URL>\n\n"
+ },
+ {
+ "type": "code",
+ "source": "import json\nimport os\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_rl_plugin.schema import RlJobInput\n\n\ndef max_wait_time_checker(seconds: int, label: str = \"\"):\n \"\"\"Return a check() that raises TimeoutError once `seconds` have elapsed.\"\"\"\n start = time.time()\n\n def check():\n if time.time() - start > seconds:\n raise TimeoutError(f\"{label} took longer than {seconds} seconds\")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nsdk = NeMoPlatform(base_url=NMP_BASE_URL, workspace=\"default\")",
+ "language": "python",
+ "source_html": "import json\nimport os\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_rl_plugin.schema import RlJobInput\n\n\ndef max_wait_time_checker(seconds: int, label: str = ""):\n """Return a check() that raises TimeoutError once `seconds` have elapsed."""\n start = time.time()\n\n def check():\n if time.time() - start > seconds:\n raise TimeoutError(f"{label} took longer than {seconds} seconds")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nsdk = NeMoPlatform(base_url=NMP_BASE_URL, workspace="default")\n"
+ },
+ {
+ "type": "markdown",
+ "source": "### 2. Prepare the Preference Dataset\n\nDPO trains on **preference data**. The `rl` backend takes a **single** dataset fileset that holds both `training.jsonl` and `validation.jsonl`, and auto-detects the row schema from the first line. Three preference formats are supported (see the platform's `BinaryPreferenceDatasetItemSchema` / `HelpSteer3DatasetItemSchema` / `Tulu3PreferenceDatasetItemSchema`):",
+ "source_html": "DPO trains on preference data. The rl backend takes a single dataset fileset that holds both training.jsonl and validation.jsonl, and auto-detects the row schema from the first line. Three preference formats are supported (see the platform's BinaryPreferenceDatasetItemSchema / HelpSteer3DatasetItemSchema / Tulu3PreferenceDatasetItemSchema):
Simple prompt / chosen / rejected (the prompt may be a string or a list of chat messages):
{"prompt": "What is the capital of France?", "chosen": "The capital of France is Paris.", "rejected": "I'm not sure."}\n\n"
+ },
+ {
+ "type": "markdown",
+ "source": "#### HelpSteer3 Format (used here)\n\nA conversation `context` (string or chat messages), two candidate `response1` / `response2`, and a signed `overall_preference` in -3..3 — **negative** means response 1 is preferred, **positive** means response 2, **0** is a tie. This is the **raw** schema of `nvidia/HelpSteer3`, so no conversion is needed:\n\n```json\n{\"context\": [{\"role\": \"user\", \"content\": \"Explain how to use git rebase\"}], \"response1\": \"...\", \"response2\": \"...\", \"overall_preference\": -2}\n```",
+ "source_html": "A conversation context (string or chat messages), two candidate response1 / response2, and a signed overall_preference in -3..3 — negative means response 1 is preferred, positive means response 2, 0 is a tie. This is the raw schema of nvidia/HelpSteer3, so no conversion is needed:
{"context": [{"role": "user", "content": "Explain how to use git rebase"}], "response1": "...", "response2": "...", "overall_preference": -2}\n\n"
+ },
+ {
+ "type": "markdown",
+ "source": "#### Tulu3 Preference Format\n\nFull chat conversations for both the chosen and rejected branches (each a list of messages ending with the assistant turn):\n\n```json\n{\"chosen\": [{\"role\": \"user\", \"content\": \"...\"}, {\"role\": \"assistant\", \"content\": \"preferred\"}], \"rejected\": [{\"role\": \"user\", \"content\": \"...\"}, {\"role\": \"assistant\", \"content\": \"dispreferred\"}]}\n```",
+ "source_html": "Full chat conversations for both the chosen and rejected branches (each a list of messages ending with the assistant turn):
\n{"chosen": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "preferred"}], "rejected": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "dispreferred"}]}\n\n"
+ },
+ {
+ "type": "markdown",
+ "source": "#### Download nvidia/HelpSteer3\n\nWe use [nvidia/HelpSteer3](https://huggingface.co/datasets/nvidia/HelpSteer3) (the `preference` subset), NVIDIA's open preference dataset. It ships native `train` and `validation` splits and matches the HelpSteer3 schema above, so we upload the rows **as-is** — the platform's `HelpSteer3Dataset` loader handles the `overall_preference` semantics (including ties) at training time.",
+ "source_html": "We use nvidia/HelpSteer3 (the preference subset), NVIDIA's open preference dataset. It ships native train and validation splits and matches the HelpSteer3 schema above, so we upload the rows as-is — the platform's HelpSteer3Dataset loader handles the overall_preference semantics (including ties) at training time.
Upload both JSONL files to a single FileSet so the DPO job can read them.
\n" + }, + { + "type": "code", + "source": "try:\n sdk.files.filesets.create(workspace=\"default\", name=DATASET_NAME, description=\"DPO preference data\")\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\nsdk.files.upload(local_path=DATASET_PATH, remote_path=\"\", fileset=DATASET_NAME, workspace=\"default\")\n\nprint(\"Preference data:\")\nprint(json.dumps([f.model_dump() for f in sdk.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2, default=str))", + "language": "python", + "source_html": "try:\n sdk.files.filesets.create(workspace="default", name=DATASET_NAME, description="DPO preference data")\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\nsdk.files.upload(local_path=DATASET_PATH, remote_path="", fileset=DATASET_NAME, workspace="default")\n\nprint("Preference data:")\nprint(json.dumps([f.model_dump() for f in sdk.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2, default=str))\n" + }, + { + "type": "markdown", + "source": "### 4. Secrets Setup\n\nThe base model (`meta-llama/Llama-3.2-1B-Instruct`) is gated, so store your Hugging Face token as a platform secret named `hf-token` and reference it on the model fileset.", + "source_html": "The base model (meta-llama/Llama-3.2-1B-Instruct) is gated, so store your Hugging Face token as a platform secret named hf-token and reference it on the model fileset.
DPO starts from an instruction-tuned base model. The model entity's spec is inferred asynchronously after creation.
\n" + }, + { + "type": "code", + "source": "HF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-instruct\"\n\nstorage = HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n token_secret=hf_secret.name,\n)\n\ntry:\n base_model_fs = sdk.files.filesets.create(\n workspace=\"default\", name=MODEL_NAME, description=\"Llama 3.2 1B Instruct base model\", storage=storage\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n base_model_fs = sdk.files.filesets.retrieve(workspace=\"default\", name=MODEL_NAME)\n print(\"Base model fileset already exists.\")\n\ntry:\n base_model = sdk.models.create(workspace=\"default\", name=MODEL_NAME, fileset=f\"default/{MODEL_NAME}\")\nexcept ConflictError:\n base_model = sdk.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\nprint(f\"Base model fileset: fileset://default/{base_model.name}\")\n\n# Wait for the ModelSpec to be inferred from the checkpoint.\ncheck = max_wait_time_checker(600, \"Model spec\")\nwhile not base_model.spec:\n check()\n time.sleep(10)\n base_model = sdk.models.retrieve(workspace=\"default\", name=MODEL_NAME)\nprint(\"Model spec ready\")", + "language": "python", + "source_html": "HF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-instruct"\n\nstorage = HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n token_secret=hf_secret.name,\n)\n\ntry:\n base_model_fs = sdk.files.filesets.create(\n workspace="default", name=MODEL_NAME, description="Llama 3.2 1B Instruct base model", storage=storage\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n base_model_fs = sdk.files.filesets.retrieve(workspace="default", name=MODEL_NAME)\n print("Base model fileset already exists.")\n\ntry:\n base_model = sdk.models.create(workspace="default", name=MODEL_NAME, fileset=f"default/{MODEL_NAME}")\nexcept ConflictError:\n base_model = sdk.models.retrieve(workspace="default", name=MODEL_NAME)\n\nprint(f"Base model fileset: fileset://default/{base_model.name}")\n\n# Wait for the ModelSpec to be inferred from the checkpoint.\ncheck = max_wait_time_checker(600, "Model spec")\nwhile not base_model.spec:\n check()\n time.sleep(10)\n base_model = sdk.models.retrieve(workspace="default", name=MODEL_NAME)\nprint("Model spec ready")\n" + }, + { + "type": "markdown", + "source": "### 6. Create the DPO Customization Job\n\nSubmit a DPO job to the `rl` backend with `RlJobInput`. Note the DPO-specific shape:\n\n- `model` is a string ref to the model entity; `dataset` is a **single** string ref to the preference fileset (holding both files).\n- The training method is `{\"type\": \"dpo\", ...}` — full-weight, no `finetuning_type`/LoRA.\n- `ref_policy_kl_penalty` is **β** (DPO paper): how strongly the policy stays tied to the reference model.\n- `rl` auto-generates the job id (`rl-Submit a DPO job to the rl backend with RlJobInput. Note the DPO-specific shape:
model is a string ref to the model entity; dataset is a single string ref to the preference fileset (holding both files).{"type": "dpo", ...} — full-weight, no finetuning_type/LoRA.ref_policy_kl_penalty is β (DPO paper): how strongly the policy stays tied to the reference model.rl auto-generates the job id (rl-<hex>); read it back from the response.Other configurable knobs: optimizer_type, adam_eps, activation_checkpointing, keep_top_k, val_at_end, preference_loss_weight, sft_loss_weight. Run nemo customization rl explain for the live schema.
The DPO job runs four steps: download -> dpo-training (Ray) -> upload -> model-entity. We poll the top-level job status and surface the training step's progress.
\n" + }, + { + "type": "code", + "source": "from IPython.display import clear_output\n\ncheck = max_wait_time_checker(7200, \"DPO job\")\nwhile True:\n check()\n status = sdk.jobs.get_status(name=job.job.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Job Status: {status.status}\")\n\n step = max_steps = phase = None\n for job_step in status.steps or []:\n if job_step.name == \"dpo-training\":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps, phase = d.get(\"step\"), d.get(\"max_steps\"), d.get(\"phase\")\n break\n break\n if step is not None and max_steps:\n print(f\"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)\")\n if phase:\n print(f\"Phase: {phase}\")\n\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished: {status.status}\")\n break\n time.sleep(15)\n\nassert status.status == \"completed\"", + "language": "python", + "source_html": "from IPython.display import clear_output\n\ncheck = max_wait_time_checker(7200, "DPO job")\nwhile True:\n check()\n status = sdk.jobs.get_status(name=job.job.name, workspace="default")\n clear_output(wait=True)\n print(f"Job Status: {status.status}")\n\n step = max_steps = phase = None\n for job_step in status.steps or []:\n if job_step.name == "dpo-training":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps, phase = d.get("step"), d.get("max_steps"), d.get("phase")\n break\n break\n if step is not None and max_steps:\n print(f"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)")\n if phase:\n print(f"Phase: {phase}")\n\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished: {status.status}")\n break\n time.sleep(15)\n\nassert status.status == "completed"\n" + }, + { + "type": "markdown", + "source": "**Interpreting DPO training metrics** (in `status_details.metrics`):\n\n- **`loss`** — the DPO loss; should trend down as the policy learns to separate chosen from rejected.\n- **Reward margin** (chosen minus rejected reward) — should trend **up**: the model increasingly prefers chosen responses.\n- **Validation `loss`** — watch for divergence from training loss (overfitting). Raise `ref_policy_kl_penalty` (β) or add `sft_loss_weight` if the policy drifts too far from the reference.", + "source_html": "Interpreting DPO training metrics (in status_details.metrics):
loss — the DPO loss; should trend down as the policy learns to separate chosen from rejected.loss — watch for divergence from training loss (overfitting). Raise ref_policy_kl_penalty (β) or add sft_loss_weight if the policy drifts too far from the reference.DPO produces a full-weight model entity (not an adapter). Confirm it was registered.
\n" + }, + { + "type": "code", + "source": "model_entity = sdk.models.retrieve(workspace=\"default\", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))", + "language": "python", + "source_html": "model_entity = sdk.models.retrieve(workspace="default", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))\n" + }, + { + "type": "markdown", + "source": "### 9. Deploy and Evaluate (optional)\n\nThe DPO output is a full model, so it deploys like any full-weight checkpoint (see the [Full SFT](/documentation/customizer-reference/tutorials/sft-customization-job) tutorial for details). We deploy with vLLM and send a chat completion.", + "source_html": "The DPO output is a full model, so it deploys like any full-weight checkpoint (see the Full SFT tutorial for details). We deploy with vLLM and send a chat completion.
\n" + }, + { + "type": "code", + "source": "deploy_suffix = uuid.uuid4().hex[:8]\nDEPLOYMENT_CONFIG_NAME = f\"dpo-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"dpo-deployment-{deploy_suffix}\"\n\ndeployment_config = sdk.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\"model_namespace\": \"default\", \"model_name\": OUTPUT_NAME},\n executor_config={\"gpu\": 1, \"image_name\": \"vllm/vllm-openai\", \"image_tag\": \"v0.22.1\"},\n)\n\ndeployment = sdk.inference.deployments.create(\n workspace=\"default\", name=DEPLOYMENT_NAME, config=deployment_config.name\n)\nprint(f\"Deployment name: {deployment.name}\")", + "language": "python", + "source_html": "deploy_suffix = uuid.uuid4().hex[:8]\nDEPLOYMENT_CONFIG_NAME = f"dpo-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"dpo-deployment-{deploy_suffix}"\n\ndeployment_config = sdk.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={"model_namespace": "default", "model_name": OUTPUT_NAME},\n executor_config={"gpu": 1, "image_name": "vllm/vllm-openai", "image_tag": "v0.22.1"},\n)\n\ndeployment = sdk.inference.deployments.create(\n workspace="default", name=DEPLOYMENT_NAME, config=deployment_config.name\n)\nprint(f"Deployment name: {deployment.name}")\n" + }, + { + "type": "code", + "source": "check = max_wait_time_checker(1800, \"Deployment\")\nwhile True:\n check()\n deployment_status = sdk.inference.deployments.retrieve(name=deployment.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Deployment status: {deployment_status.status}\")\n deployment_state = str(deployment_status.status).lower()\n if deployment_state in (\"ready\", \"running\"):\n if not sdk.models.wait_for_gateway(deployment.name, workspace=\"default\", timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n if deployment_state in (\"failed\", \"error\", \"terminated\", \"lost\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n time.sleep(15)", + "language": "python", + "source_html": "check = max_wait_time_checker(1800, "Deployment")\nwhile True:\n check()\n deployment_status = sdk.inference.deployments.retrieve(name=deployment.name, workspace="default")\n clear_output(wait=True)\n print(f"Deployment status: {deployment_status.status}")\n deployment_state = str(deployment_status.status).lower()\n if deployment_state in ("ready", "running"):\n if not sdk.models.wait_for_gateway(deployment.name, workspace="default", timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n if deployment_state in ("failed", "error", "terminated", "lost"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n time.sleep(15)\n" + }, + { + "type": "code", + "source": "messages = [\n {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n {\"role\": \"user\", \"content\": \"Write a short, friendly email to a colleague asking to reschedule our meeting to Thursday.\"},\n]\n\nresponse = sdk.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\"model\": f\"default/{OUTPUT_NAME}\", \"messages\": messages, \"temperature\": 0.7, \"max_tokens\": 256},\n)\nprint(\"Model output:\\n\")\nprint(response[\"choices\"][0][\"message\"][\"content\"])", + "language": "python", + "source_html": "messages = [\n {"role": "system", "content": "You are a helpful assistant."},\n {"role": "user", "content": "Write a short, friendly email to a colleague asking to reschedule our meeting to Thursday."},\n]\n\nresponse = sdk.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={"model": f"default/{OUTPUT_NAME}", "messages": messages, "temperature": 0.7, "max_tokens": 256},\n)\nprint("Model output:\\n")\nprint(response["choices"][0]["message"]["content"])\n" + }, + { + "type": "markdown", + "source": "## Conclusion\n\nYou aligned a base model with **DPO** on the NeMo Platform using the `rl` backend:\n\n- Uploaded a HelpSteer3 preference dataset **as-is** (the platform detects the schema natively).\n- Submitted a full-weight DPO job that ran on a Ray cluster via the Kubernetes executor.\n- Registered the output as a full model entity and (optionally) deployed it for inference.\n\n**Next steps:** tune the alignment strength with `ref_policy_kl_penalty` (β), add `sft_loss_weight` to anchor the policy to the chosen responses, enable `activation_checkpointing` for memory headroom, or scale up with `parallelism`. See the [Training Configuration](/documentation/customizer-reference/manage-customization-jobs/training-configuration) reference for the full hyperparameter set.", + "source_html": "You aligned a base model with DPO on the NeMo Platform using the rl backend:
Next steps: tune the alignment strength with ref_policy_kl_penalty (β), add sft_loss_weight to anchor the policy to the chosen responses, enable activation_checkpointing for memory headroom, or scale up with parallelism. See the Training Configuration reference for the full hyperparameter set.
Learn how to use the NeMo Platform to align a model with DPO (Direct Preference Optimization) on a preference dataset. For each prompt, DPO trains on a chosen (preferred) and a rejected response so the model prefers the chosen style — no separate reward model required.
\nThis tutorial uses the rl customization backend (powered by NVIDIA NeMo-RL), which runs DPO on a Ray cluster. Unlike the SFT and LoRA tutorials (Docker GPU jobs), rl requires a Kubernetes-backed NeMo Platform. DPO here is full-weight (no LoRA/adapter); the output is a full model entity.
Time to complete: approximately 45-60 minutes. Job duration increases with model and dataset size.
\n" + }, + { + "type": "markdown", + "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](/documentation/get-started)** to install the NeMo Platform and Python SDK.\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root).\n3. **Installed the `datasets` package**: `pip install datasets`.\n4. **A platform configured with `platform.runtime: kubernetes`.** The `rl` (DPO) backend provisions a Ray cluster and has **no local Docker fallback** — `submit` fails fast on a Docker-runtime platform. Multi-node jobs (`parallelism.num_nodes > 1`) additionally require the platform-side `NMP_RL_MULTINODE_SHARED_STORAGE_PATH`.\n5. **A Hugging Face token** with access to the gated base model (this tutorial uses `meta-llama/Llama-3.2-1B-Instruct`). Export it as `HF_TOKEN`.\n6. **At least one GPU with CUDA 13+** and a GPU execution profile (`nemo jobs list-execution-profiles`).", + "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root).datasets package: pip install datasets.platform.runtime: kubernetes. The rl (DPO) backend provisions a Ray cluster and has no local Docker fallback — submit fails fast on a Docker-runtime platform. Multi-node jobs (parallelism.num_nodes > 1) additionally require the platform-side NMP_RL_MULTINODE_SHARED_STORAGE_PATH.meta-llama/Llama-3.2-1B-Instruct). Export it as HF_TOKEN.nemo jobs list-execution-profiles).The SDK needs your NeMo Platform server URL. By default http://localhost:8080 is used; set NMP_BASE_URL to override:
export NMP_BASE_URL=<YOUR_NMP_BASE_URL>\n\n"
+ },
+ {
+ "type": "code",
+ "source": "import json\nimport os\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_rl_plugin.schema import RlJobInput\n\n\ndef max_wait_time_checker(seconds: int, label: str = \"\"):\n \"\"\"Return a check() that raises TimeoutError once `seconds` have elapsed.\"\"\"\n start = time.time()\n\n def check():\n if time.time() - start > seconds:\n raise TimeoutError(f\"{label} took longer than {seconds} seconds\")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nsdk = NeMoPlatform(base_url=NMP_BASE_URL, workspace=\"default\")",
+ "language": "python",
+ "source_html": "import json\nimport os\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_rl_plugin.schema import RlJobInput\n\n\ndef max_wait_time_checker(seconds: int, label: str = ""):\n """Return a check() that raises TimeoutError once `seconds` have elapsed."""\n start = time.time()\n\n def check():\n if time.time() - start > seconds:\n raise TimeoutError(f"{label} took longer than {seconds} seconds")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nsdk = NeMoPlatform(base_url=NMP_BASE_URL, workspace="default")\n"
+ },
+ {
+ "type": "markdown",
+ "source": "### 2. Prepare the Preference Dataset\n\nDPO trains on **preference data**. The `rl` backend takes a **single** dataset fileset that holds both `training.jsonl` and `validation.jsonl`, and auto-detects the row schema from the first line. Three preference formats are supported (see the platform's `BinaryPreferenceDatasetItemSchema` / `HelpSteer3DatasetItemSchema` / `Tulu3PreferenceDatasetItemSchema`):",
+ "source_html": "DPO trains on preference data. The rl backend takes a single dataset fileset that holds both training.jsonl and validation.jsonl, and auto-detects the row schema from the first line. Three preference formats are supported (see the platform's BinaryPreferenceDatasetItemSchema / HelpSteer3DatasetItemSchema / Tulu3PreferenceDatasetItemSchema):
Simple prompt / chosen / rejected (the prompt may be a string or a list of chat messages):
{"prompt": "What is the capital of France?", "chosen": "The capital of France is Paris.", "rejected": "I'm not sure."}\n\n"
+ },
+ {
+ "type": "markdown",
+ "source": "#### HelpSteer3 Format (used here)\n\nA conversation `context` (string or chat messages), two candidate `response1` / `response2`, and a signed `overall_preference` in -3..3 — **negative** means response 1 is preferred, **positive** means response 2, **0** is a tie. This is the **raw** schema of `nvidia/HelpSteer3`, so no conversion is needed:\n\n```json\n{\"context\": [{\"role\": \"user\", \"content\": \"Explain how to use git rebase\"}], \"response1\": \"...\", \"response2\": \"...\", \"overall_preference\": -2}\n```",
+ "source_html": "A conversation context (string or chat messages), two candidate response1 / response2, and a signed overall_preference in -3..3 — negative means response 1 is preferred, positive means response 2, 0 is a tie. This is the raw schema of nvidia/HelpSteer3, so no conversion is needed:
{"context": [{"role": "user", "content": "Explain how to use git rebase"}], "response1": "...", "response2": "...", "overall_preference": -2}\n\n"
+ },
+ {
+ "type": "markdown",
+ "source": "#### Tulu3 Preference Format\n\nFull chat conversations for both the chosen and rejected branches (each a list of messages ending with the assistant turn):\n\n```json\n{\"chosen\": [{\"role\": \"user\", \"content\": \"...\"}, {\"role\": \"assistant\", \"content\": \"preferred\"}], \"rejected\": [{\"role\": \"user\", \"content\": \"...\"}, {\"role\": \"assistant\", \"content\": \"dispreferred\"}]}\n```",
+ "source_html": "Full chat conversations for both the chosen and rejected branches (each a list of messages ending with the assistant turn):
\n{"chosen": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "preferred"}], "rejected": [{"role": "user", "content": "..."}, {"role": "assistant", "content": "dispreferred"}]}\n\n"
+ },
+ {
+ "type": "markdown",
+ "source": "#### Download nvidia/HelpSteer3\n\nWe use [nvidia/HelpSteer3](https://huggingface.co/datasets/nvidia/HelpSteer3) (the `preference` subset), NVIDIA's open preference dataset. It ships native `train` and `validation` splits and matches the HelpSteer3 schema above, so we upload the rows **as-is** — the platform's `HelpSteer3Dataset` loader handles the `overall_preference` semantics (including ties) at training time.",
+ "source_html": "We use nvidia/HelpSteer3 (the preference subset), NVIDIA's open preference dataset. It ships native train and validation splits and matches the HelpSteer3 schema above, so we upload the rows as-is — the platform's HelpSteer3Dataset loader handles the overall_preference semantics (including ties) at training time.
Upload both JSONL files to a single FileSet so the DPO job can read them.
\n" + }, + { + "type": "code", + "source": "try:\n sdk.files.filesets.create(workspace=\"default\", name=DATASET_NAME, description=\"DPO preference data\")\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\nsdk.files.upload(local_path=DATASET_PATH, remote_path=\"\", fileset=DATASET_NAME, workspace=\"default\")\n\nprint(\"Preference data:\")\nprint(json.dumps([f.model_dump() for f in sdk.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2, default=str))", + "language": "python", + "source_html": "try:\n sdk.files.filesets.create(workspace="default", name=DATASET_NAME, description="DPO preference data")\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\nsdk.files.upload(local_path=DATASET_PATH, remote_path="", fileset=DATASET_NAME, workspace="default")\n\nprint("Preference data:")\nprint(json.dumps([f.model_dump() for f in sdk.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2, default=str))\n" + }, + { + "type": "markdown", + "source": "### 4. Secrets Setup\n\nThe base model (`meta-llama/Llama-3.2-1B-Instruct`) is gated, so store your Hugging Face token as a platform secret named `hf-token` and reference it on the model fileset.", + "source_html": "The base model (meta-llama/Llama-3.2-1B-Instruct) is gated, so store your Hugging Face token as a platform secret named hf-token and reference it on the model fileset.
DPO starts from an instruction-tuned base model. The model entity's spec is inferred asynchronously after creation.
\n" + }, + { + "type": "code", + "source": "HF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-instruct\"\n\nstorage = HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n token_secret=hf_secret.name,\n)\n\ntry:\n base_model_fs = sdk.files.filesets.create(\n workspace=\"default\", name=MODEL_NAME, description=\"Llama 3.2 1B Instruct base model\", storage=storage\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n base_model_fs = sdk.files.filesets.retrieve(workspace=\"default\", name=MODEL_NAME)\n print(\"Base model fileset already exists.\")\n\ntry:\n base_model = sdk.models.create(workspace=\"default\", name=MODEL_NAME, fileset=f\"default/{MODEL_NAME}\")\nexcept ConflictError:\n base_model = sdk.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\nprint(f\"Base model fileset: fileset://default/{base_model.name}\")\n\n# Wait for the ModelSpec to be inferred from the checkpoint.\ncheck = max_wait_time_checker(600, \"Model spec\")\nwhile not base_model.spec:\n check()\n time.sleep(10)\n base_model = sdk.models.retrieve(workspace=\"default\", name=MODEL_NAME)\nprint(\"Model spec ready\")", + "language": "python", + "source_html": "HF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-instruct"\n\nstorage = HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n token_secret=hf_secret.name,\n)\n\ntry:\n base_model_fs = sdk.files.filesets.create(\n workspace="default", name=MODEL_NAME, description="Llama 3.2 1B Instruct base model", storage=storage\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n base_model_fs = sdk.files.filesets.retrieve(workspace="default", name=MODEL_NAME)\n print("Base model fileset already exists.")\n\ntry:\n base_model = sdk.models.create(workspace="default", name=MODEL_NAME, fileset=f"default/{MODEL_NAME}")\nexcept ConflictError:\n base_model = sdk.models.retrieve(workspace="default", name=MODEL_NAME)\n\nprint(f"Base model fileset: fileset://default/{base_model.name}")\n\n# Wait for the ModelSpec to be inferred from the checkpoint.\ncheck = max_wait_time_checker(600, "Model spec")\nwhile not base_model.spec:\n check()\n time.sleep(10)\n base_model = sdk.models.retrieve(workspace="default", name=MODEL_NAME)\nprint("Model spec ready")\n" + }, + { + "type": "markdown", + "source": "### 6. Create the DPO Customization Job\n\nSubmit a DPO job to the `rl` backend with `RlJobInput`. Note the DPO-specific shape:\n\n- `model` is a string ref to the model entity; `dataset` is a **single** string ref to the preference fileset (holding both files).\n- The training method is `{\"type\": \"dpo\", ...}` — full-weight, no `finetuning_type`/LoRA.\n- `ref_policy_kl_penalty` is **β** (DPO paper): how strongly the policy stays tied to the reference model.\n- `rl` auto-generates the job id (`rl-Submit a DPO job to the rl backend with RlJobInput. Note the DPO-specific shape:
model is a string ref to the model entity; dataset is a single string ref to the preference fileset (holding both files).{"type": "dpo", ...} — full-weight, no finetuning_type/LoRA.ref_policy_kl_penalty is β (DPO paper): how strongly the policy stays tied to the reference model.rl auto-generates the job id (rl-<hex>); read it back from the response.Other configurable knobs: optimizer_type, adam_eps, activation_checkpointing, keep_top_k, val_at_end, preference_loss_weight, sft_loss_weight. Run nemo customization rl explain for the live schema.
The DPO job runs four steps: download -> dpo-training (Ray) -> upload -> model-entity. We poll the top-level job status and surface the training step's progress.
\n" + }, + { + "type": "code", + "source": "from IPython.display import clear_output\n\ncheck = max_wait_time_checker(7200, \"DPO job\")\nwhile True:\n check()\n status = sdk.jobs.get_status(name=job.job.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Job Status: {status.status}\")\n\n step = max_steps = phase = None\n for job_step in status.steps or []:\n if job_step.name == \"dpo-training\":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps, phase = d.get(\"step\"), d.get(\"max_steps\"), d.get(\"phase\")\n break\n break\n if step is not None and max_steps:\n print(f\"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)\")\n if phase:\n print(f\"Phase: {phase}\")\n\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished: {status.status}\")\n break\n time.sleep(15)\n\nassert status.status == \"completed\"", + "language": "python", + "source_html": "from IPython.display import clear_output\n\ncheck = max_wait_time_checker(7200, "DPO job")\nwhile True:\n check()\n status = sdk.jobs.get_status(name=job.job.name, workspace="default")\n clear_output(wait=True)\n print(f"Job Status: {status.status}")\n\n step = max_steps = phase = None\n for job_step in status.steps or []:\n if job_step.name == "dpo-training":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps, phase = d.get("step"), d.get("max_steps"), d.get("phase")\n break\n break\n if step is not None and max_steps:\n print(f"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)")\n if phase:\n print(f"Phase: {phase}")\n\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished: {status.status}")\n break\n time.sleep(15)\n\nassert status.status == "completed"\n" + }, + { + "type": "markdown", + "source": "**Interpreting DPO training metrics** (in `status_details.metrics`):\n\n- **`loss`** — the DPO loss; should trend down as the policy learns to separate chosen from rejected.\n- **Reward margin** (chosen minus rejected reward) — should trend **up**: the model increasingly prefers chosen responses.\n- **Validation `loss`** — watch for divergence from training loss (overfitting). Raise `ref_policy_kl_penalty` (β) or add `sft_loss_weight` if the policy drifts too far from the reference.", + "source_html": "Interpreting DPO training metrics (in status_details.metrics):
loss — the DPO loss; should trend down as the policy learns to separate chosen from rejected.loss — watch for divergence from training loss (overfitting). Raise ref_policy_kl_penalty (β) or add sft_loss_weight if the policy drifts too far from the reference.DPO produces a full-weight model entity (not an adapter). Confirm it was registered.
\n" + }, + { + "type": "code", + "source": "model_entity = sdk.models.retrieve(workspace=\"default\", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))", + "language": "python", + "source_html": "model_entity = sdk.models.retrieve(workspace="default", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))\n" + }, + { + "type": "markdown", + "source": "### 9. Deploy and Evaluate (optional)\n\nThe DPO output is a full model, so it deploys like any full-weight checkpoint (see the [Full SFT](/documentation/customizer-reference/tutorials/sft-customization-job) tutorial for details). We deploy with vLLM and send a chat completion.", + "source_html": "The DPO output is a full model, so it deploys like any full-weight checkpoint (see the Full SFT tutorial for details). We deploy with vLLM and send a chat completion.
\n" + }, + { + "type": "code", + "source": "deploy_suffix = uuid.uuid4().hex[:8]\nDEPLOYMENT_CONFIG_NAME = f\"dpo-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"dpo-deployment-{deploy_suffix}\"\n\ndeployment_config = sdk.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\"model_namespace\": \"default\", \"model_name\": OUTPUT_NAME},\n executor_config={\"gpu\": 1, \"image_name\": \"vllm/vllm-openai\", \"image_tag\": \"v0.22.1\"},\n)\n\ndeployment = sdk.inference.deployments.create(\n workspace=\"default\", name=DEPLOYMENT_NAME, config=deployment_config.name\n)\nprint(f\"Deployment name: {deployment.name}\")", + "language": "python", + "source_html": "deploy_suffix = uuid.uuid4().hex[:8]\nDEPLOYMENT_CONFIG_NAME = f"dpo-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"dpo-deployment-{deploy_suffix}"\n\ndeployment_config = sdk.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={"model_namespace": "default", "model_name": OUTPUT_NAME},\n executor_config={"gpu": 1, "image_name": "vllm/vllm-openai", "image_tag": "v0.22.1"},\n)\n\ndeployment = sdk.inference.deployments.create(\n workspace="default", name=DEPLOYMENT_NAME, config=deployment_config.name\n)\nprint(f"Deployment name: {deployment.name}")\n" + }, + { + "type": "code", + "source": "check = max_wait_time_checker(1800, \"Deployment\")\nwhile True:\n check()\n deployment_status = sdk.inference.deployments.retrieve(name=deployment.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Deployment status: {deployment_status.status}\")\n deployment_state = str(deployment_status.status).lower()\n if deployment_state in (\"ready\", \"running\"):\n if not sdk.models.wait_for_gateway(deployment.name, workspace=\"default\", timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n if deployment_state in (\"failed\", \"error\", \"terminated\", \"lost\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n time.sleep(15)", + "language": "python", + "source_html": "check = max_wait_time_checker(1800, "Deployment")\nwhile True:\n check()\n deployment_status = sdk.inference.deployments.retrieve(name=deployment.name, workspace="default")\n clear_output(wait=True)\n print(f"Deployment status: {deployment_status.status}")\n deployment_state = str(deployment_status.status).lower()\n if deployment_state in ("ready", "running"):\n if not sdk.models.wait_for_gateway(deployment.name, workspace="default", timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n if deployment_state in ("failed", "error", "terminated", "lost"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n time.sleep(15)\n" + }, + { + "type": "code", + "source": "messages = [\n {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n {\"role\": \"user\", \"content\": \"Write a short, friendly email to a colleague asking to reschedule our meeting to Thursday.\"},\n]\n\nresponse = sdk.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\"model\": f\"default/{OUTPUT_NAME}\", \"messages\": messages, \"temperature\": 0.7, \"max_tokens\": 256},\n)\nprint(\"Model output:\\n\")\nprint(response[\"choices\"][0][\"message\"][\"content\"])", + "language": "python", + "source_html": "messages = [\n {"role": "system", "content": "You are a helpful assistant."},\n {"role": "user", "content": "Write a short, friendly email to a colleague asking to reschedule our meeting to Thursday."},\n]\n\nresponse = sdk.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={"model": f"default/{OUTPUT_NAME}", "messages": messages, "temperature": 0.7, "max_tokens": 256},\n)\nprint("Model output:\\n")\nprint(response["choices"][0]["message"]["content"])\n" + }, + { + "type": "markdown", + "source": "## Conclusion\n\nYou aligned a base model with **DPO** on the NeMo Platform using the `rl` backend:\n\n- Uploaded a HelpSteer3 preference dataset **as-is** (the platform detects the schema natively).\n- Submitted a full-weight DPO job that ran on a Ray cluster via the Kubernetes executor.\n- Registered the output as a full model entity and (optionally) deployed it for inference.\n\n**Next steps:** tune the alignment strength with `ref_policy_kl_penalty` (β), add `sft_loss_weight` to anchor the policy to the chosen responses, enable `activation_checkpointing` for memory headroom, or scale up with `parallelism`. See the [Training Configuration](/documentation/customizer-reference/manage-customization-jobs/training-configuration) reference for the full hyperparameter set.", + "source_html": "You aligned a base model with DPO on the NeMo Platform using the rl backend:
Next steps: tune the alignment strength with ref_policy_kl_penalty (β), add sft_loss_weight to anchor the policy to the chosen responses, enable activation_checkpointing for memory headroom, or scale up with parallelism. See the Training Configuration reference for the full hyperparameter set.
Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)The SPECTER dataset and the tutorial's base model are public and do not require a Hugging Face token. If you substitute a gated or private model, provide a token with read access.
\n" }, { "type": "markdown", @@ -64,8 +64,8 @@ }, { "type": "markdown", - "source": "### 3. Prepare Dataset\n\nUse the [SPECTER dataset](https://huggingface.co/datasets/embedding-data/SPECTER) from HuggingFace, a collection of scientific paper triplets where papers that cite each other are considered related.\n\n**Dataset structure:**\n- ~684K scientific paper triplets (this tutorial uses 10%)\n- Each triplet: (query paper, positive/related paper, negative/unrelated paper)\n- Papers that cite each other are marked as \"related\"\n\nIn this tutorial the following dataset directory structure will be used:\n```\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n```", - "source_html": "Use the SPECTER dataset from HuggingFace, a collection of scientific paper triplets where papers that cite each other are considered related.
\nDataset structure:
\nIn this tutorial the following dataset directory structure will be used:
\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n\n"
+ "source": "### 3. Prepare Dataset\n\nUse the [SPECTER dataset](https://huggingface.co/datasets/embedding-data/SPECTER) from Hugging Face, a collection of scientific paper triplets where papers that cite each other are considered related.\n\n**Dataset structure:**\n- ~684K scientific paper triplets (this tutorial uses 10%)\n- Each triplet: (query paper, positive/related paper, negative/unrelated paper)\n- Papers that cite each other are marked as \"related\"\n\nIn this tutorial the following dataset directory structure will be used:\n```\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n```",
+ "source_html": "Use the SPECTER dataset from Hugging Face, a collection of scientific paper triplets where papers that cite each other are considered related.
\nDataset structure:
\nIn this tutorial the following dataset directory structure will be used:
\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n\n"
},
{
"type": "markdown",
@@ -74,9 +74,9 @@
},
{
"type": "code",
- "source": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# HuggingFace token for dataset access\nHF_TOKEN = os.environ.get(\"HF_TOKEN\")\nif not HF_TOKEN:\n raise ValueError(\"HF_TOKEN environment variable is required. Get one at https://huggingface.co/settings/tokens\")\nos.environ[\"HF_TOKEN\"] = HF_TOKEN\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path(\"embedding-dataset\").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint(\"Downloading SPECTER dataset...\")\ndata = load_dataset(\"embedding-data/SPECTER\")[\"train\"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint(\"Splitting into train/validation...\")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits[\"train\"]\nvalidation_data = splits[\"test\"]\n\n# Convert to triplet JSONL format\nprint(\"Saving to JSONL...\")\nfor name, dataset in [(\"training\", train_data), (\"validation\", validation_data)]:\n with open(f\"{DATASET_PATH}/{name}.jsonl\", \"w\") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n \"query\": row[\"set\"][0],\n \"pos_doc\": row[\"set\"][1],\n \"neg_doc\": [row[\"set\"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + \"\\n\")\n\nprint(f\"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples\")\nprint(f\"\\nExample triplet:\")\nprint(f\" Query: {train_data[0]['set'][0][:100]}...\")\nprint(f\" Positive: {train_data[0]['set'][1][:100]}...\")\nprint(f\" Negative: {train_data[0]['set'][2][:100]}...\")",
+ "source": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path(\"embedding-dataset\").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint(\"Downloading SPECTER dataset...\")\ndata = load_dataset(\"embedding-data/SPECTER\")[\"train\"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint(\"Splitting into train/validation...\")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits[\"train\"]\nvalidation_data = splits[\"test\"]\n\n# Convert to triplet JSONL format\nprint(\"Saving to JSONL...\")\nfor name, dataset in [(\"training\", train_data), (\"validation\", validation_data)]:\n with open(f\"{DATASET_PATH}/{name}.jsonl\", \"w\") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n \"query\": row[\"set\"][0],\n \"pos_doc\": row[\"set\"][1],\n \"neg_doc\": [row[\"set\"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + \"\\n\")\n\nprint(f\"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples\")\nprint(f\"\\nExample triplet:\")\nprint(f\" Query: {train_data[0]['set'][0][:100]}...\")\nprint(f\" Positive: {train_data[0]['set'][1][:100]}...\")\nprint(f\" Negative: {train_data[0]['set'][2][:100]}...\")",
"language": "python",
- "source_html": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# HuggingFace token for dataset access\nHF_TOKEN = os.environ.get("HF_TOKEN")\nif not HF_TOKEN:\n raise ValueError("HF_TOKEN environment variable is required. Get one at https://huggingface.co/settings/tokens")\nos.environ["HF_TOKEN"] = HF_TOKEN\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path("embedding-dataset").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint("Downloading SPECTER dataset...")\ndata = load_dataset("embedding-data/SPECTER")["train"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint("Splitting into train/validation...")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits["train"]\nvalidation_data = splits["test"]\n\n# Convert to triplet JSONL format\nprint("Saving to JSONL...")\nfor name, dataset in [("training", train_data), ("validation", validation_data)]:\n with open(f"{DATASET_PATH}/{name}.jsonl", "w") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n "query": row["set"][0],\n "pos_doc": row["set"][1],\n "neg_doc": [row["set"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + "\\n")\n\nprint(f"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples")\nprint(f"\\nExample triplet:")\nprint(f" Query: {train_data[0]['set'][0][:100]}...")\nprint(f" Positive: {train_data[0]['set'][1][:100]}...")\nprint(f" Negative: {train_data[0]['set'][2][:100]}...")\n"
+ "source_html": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path("embedding-dataset").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint("Downloading SPECTER dataset...")\ndata = load_dataset("embedding-data/SPECTER")["train"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint("Splitting into train/validation...")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits["train"]\nvalidation_data = splits["test"]\n\n# Convert to triplet JSONL format\nprint("Saving to JSONL...")\nfor name, dataset in [("training", train_data), ("validation", validation_data)]:\n with open(f"{DATASET_PATH}/{name}.jsonl", "w") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n "query": row["set"][0],\n "pos_doc": row["set"][1],\n "neg_doc": [row["set"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + "\\n")\n\nprint(f"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples")\nprint(f"\\nExample triplet:")\nprint(f" Query: {train_data[0]['set'][0][:100]}...")\nprint(f" Positive: {train_data[0]['set'][1][:100]}...")\nprint(f" Negative: {train_data[0]['set'][2][:100]}...")\n"
},
{
"type": "markdown",
@@ -91,30 +91,30 @@
},
{
"type": "markdown",
- "source": "### 6. Secrets Setup\n\nConfigure authentication for accessing base models:\n\n- **NGC models** (`ngc://` URIs): Requires NGC API key\n- **HuggingFace models** (`hf://` URIs): Requires HF token for gated/private models\n\nGet your credentials:\n- [NGC API Key](https://ngc.nvidia.com/) (Setup → Generate API Key)\n- [HuggingFace Token](https://huggingface.co/settings/tokens) (Create token with Read access)\n\n---\n\n#### Quick Setup Example\n\nThis tutorial fine-tunes [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2), an NVIDIA embedding model optimized for question-answering and retrieval tasks.",
- "source_html": "Configure authentication for accessing base models:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsGet your credentials:
\nThis tutorial fine-tunes nvidia/llama-nemotron-embed-1b-v2, an NVIDIA embedding model optimized for question-answering and retrieval tasks.
\n" + "source": "### 6. Secrets Setup\n\nConfigure authentication for accessing base models:\n\n- **NGC models** (`ngc://` URIs): Requires NGC API key\n- **Hugging Face models** (`hf://` URIs): Requires HF token for gated/private models\n\nGet your credentials:\n- [NGC API Key](https://ngc.nvidia.com/) (Setup → Generate API Key)\n- [Hugging Face Token](https://huggingface.co/settings/tokens) (Optional; needed only for a gated/private replacement model)\n\n---\n\n#### Quick Setup Example\n\nThis tutorial fine-tunes [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2), an NVIDIA embedding model optimized for question-answering and retrieval tasks.", + "source_html": "Configure authentication for accessing base models:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsGet your credentials:
\nThis tutorial fine-tunes nvidia/llama-nemotron-embed-1b-v2, an NVIDIA embedding model optimized for question-answering and retrieval tasks.
\n" }, { "type": "code", - "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif \"NGC_SECRET_NAME\" not in globals():\n NGC_SECRET_NAME = \"ngc-api-key\"\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME}\")", + "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Public Hugging Face models need no token. Create a secret only when HF_TOKEN is set.\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\") if HF_TOKEN else None\nif hf_secret:\n print(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif \"NGC_SECRET_NAME\" not in globals():\n NGC_SECRET_NAME = \"ngc-api-key\"\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME}\")", "language": "python", - "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif "NGC_SECRET_NAME" not in globals():\n NGC_SECRET_NAME = "ngc-api-key"\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME}")\n" + "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Public Hugging Face models need no token. Create a secret only when HF_TOKEN is set.\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN") if HF_TOKEN else None\nif hf_secret:\n print(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif "NGC_SECRET_NAME" not in globals():\n NGC_SECRET_NAME = "ngc-api-key"\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME}")\n" }, { "type": "markdown", - "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2) embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.", - "source_html": "Create a fileset pointing to the nvidia/llama-nemotron-embed-1b-v2 embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n" + "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2) embedding model from Hugging Face, then create a Model Entity that references this fileset. Model downloading will take place at training time.", + "source_html": "Create a fileset pointing to the nvidia/llama-nemotron-embed-1b-v2 embedding model from Hugging Face, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n" }, { "type": "code", - "source": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"nvidia/llama-nemotron-embed-1b-v2\"\nMODEL_NAME = \"nv-nemotron-embed-1b-base\"\n\n# Ensure you have a HuggingFace token secret created\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"NVIDIA Llama Nemotron Embed 1B v2 embedding model\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n # repo_id is the full model name from Hugging Face\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n # we use the secret created in the previous step\n token_secret=hf_secret.name\n )\n )\nexcept ConflictError as e:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n\nprint(f\"\\nBase model fileset: fileset://default/{base_model.name}\")\nprint(\"\\nBase model files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=\"default\").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec}\")", + "source": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"nvidia/llama-nemotron-embed-1b-v2\"\nMODEL_NAME = \"nv-nemotron-embed-1b-base\"\n\nstorage_kwargs = {\n \"type\": \"huggingface\",\n \"repo_id\": HF_REPO_ID,\n \"repo_type\": \"model\",\n}\nif hf_secret:\n storage_kwargs[\"token_secret\"] = hf_secret.name\nstorage = HuggingfaceStorageConfigParam(**storage_kwargs)\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"NVIDIA Llama Nemotron Embed 1B v2 embedding model\",\n storage=storage,\n )\nexcept ConflictError as e:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n\nprint(f\"\\nBase model fileset: fileset://default/{base_model.name}\")\nprint(\"\\nBase model files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=\"default\").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec}\")", "language": "python", - "source_html": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "nvidia/llama-nemotron-embed-1b-v2"\nMODEL_NAME = "nv-nemotron-embed-1b-base"\n\n# Ensure you have a HuggingFace token secret created\ntry:\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="NVIDIA Llama Nemotron Embed 1B v2 embedding model",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n # repo_id is the full model name from Hugging Face\n repo_id=HF_REPO_ID,\n repo_type="model",\n # we use the secret created in the previous step\n token_secret=hf_secret.name\n )\n )\nexcept ConflictError as e:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n\nprint(f"\\nBase model fileset: fileset://default/{base_model.name}")\nprint("\\nBase model files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace="default").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec}")\n" + "source_html": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "nvidia/llama-nemotron-embed-1b-v2"\nMODEL_NAME = "nv-nemotron-embed-1b-base"\n\nstorage_kwargs = {\n "type": "huggingface",\n "repo_id": HF_REPO_ID,\n "repo_type": "model",\n}\nif hf_secret:\n storage_kwargs["token_secret"] = hf_secret.name\nstorage = HuggingfaceStorageConfigParam(**storage_kwargs)\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="NVIDIA Llama Nemotron Embed 1B v2 embedding model",\n storage=storage,\n )\nexcept ConflictError as e:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n\nprint(f"\\nBase model fileset: fileset://default/{base_model.name}")\nprint("\\nBase model files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace="default").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec}")\n" }, { "type": "markdown", - "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with split `schedule`, `batch`, `optimizer`, and `parallelism` sections. Reference the model entity and dataset fileset by workspace/name (not `fileset://` URIs).\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training.training_type`**: `sft`\n- **`training.finetuning_type`**: `all_weights` for full fine-tuning, or `lora_merged` for merged LoRA\n- **`optimizer.learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch.global_batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set `finetuning_type` to `lora_merged`:\n\n```python\ntraining={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora_merged\",\n \"lora\": {\"rank\": 16, \"alpha\": 32},\n \"max_seq_length\": MAX_SEQ_LENGTH,\n}\n```", - "source_html": "Create a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nSubmit to the Automodel backend using AutomodelJobInput with split schedule, batch, optimizer, and parallelism sections. Reference the model entity and dataset fileset by workspace/name (not fileset:// URIs).
Key hyperparameters for embedding fine-tuning:
\ntraining.training_type: sfttraining.finetuning_type: all_weights for full fine-tuning, or lora_merged for merged LoRAoptimizer.learning_rate: Lower values (1e-6 to 5e-6) work well for embedding modelsbatch.global_batch_size: Larger batches improve contrastive learning (128-256 recommended)NOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set finetuning_type to lora_merged:
training={\n "training_type": "sft",\n "finetuning_type": "lora_merged",\n "lora": {"rank": 16, "alpha": 32},\n "max_seq_length": MAX_SEQ_LENGTH,\n}\n\n"
+ "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with split `schedule`, `batch`, `optimizer`, and `parallelism` sections. Reference the model entity and dataset fileset by workspace/name (not `fileset://` URIs).\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training.training_type`**: `sft`\n- **`training.finetuning_type`**: `all_weights` for full fine-tuning, or `lora_merged` for merged LoRA\n- **`optimizer.learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch.global_batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set `finetuning_type` to `lora_merged`:\n\n```python\ntraining={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora_merged\",\n \"lora\": {\"rank\": 16, \"alpha\": 32},\n \"max_seq_length\": 512,\n}\n```",
+ "source_html": "Create a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nSubmit to the Automodel backend using AutomodelJobInput with split schedule, batch, optimizer, and parallelism sections. Reference the model entity and dataset fileset by workspace/name (not fileset:// URIs).
Key hyperparameters for embedding fine-tuning:
\ntraining.training_type: sfttraining.finetuning_type: all_weights for full fine-tuning, or lora_merged for merged LoRAoptimizer.learning_rate: Lower values (1e-6 to 5e-6) work well for embedding modelsbatch.global_batch_size: Larger batches improve contrastive learning (128-256 recommended)NOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set finetuning_type to lora_merged:
training={\n "training_type": "sft",\n "finetuning_type": "lora_merged",\n "lora": {"rank": 16, "alpha": 32},\n "max_seq_length": 512,\n}\n\n"
},
{
"type": "code",
@@ -129,9 +129,9 @@
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
+ "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)\n\nif status.status != \"completed\":\n raise RuntimeError(f\"Training job finished with status: {status.status}\")",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n\nif status.status != "completed":\n raise RuntimeError(f"Training job finished with status: {status.status}")\n"
},
{
"type": "markdown",
@@ -179,8 +179,8 @@
},
{
"type": "markdown",
- "source": "### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world queries from your domain\n- Compare retrieval rankings before and after fine-tuning\n- Check that semantically similar items rank higher than keyword matches\n\n**What to look for:**\n- ✅ Relevant documents consistently rank in top positions\n- ✅ Keyword traps (like \"Random Forest\" vs \"Random Fields\") are handled correctly\n- ✅ Domain-specific terminology is understood\n- ❌ Unrelated documents with matching keywords do not rank high\n\n**Benchmark Evaluation**\n\nFor systematic evaluation, use the NeMo Evaluator service with retrieval benchmarks like SciDocs, BEIR, or MTEB. Refer to the [Evaluator documentation](../../evaluator/index.md) for details.\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n**Embedding-Specific Recommendations:**\n\n| Parameter | Recommended | Notes |\n|-----------|-------------|-------|\n| `learning_rate` | 1e-6 to 5e-6 | Lower than standard SFT |\n| `batch_size` | 128-256 | Larger batches improve contrastive learning |\n| `max_seq_length` | 512 | Typical for embedding models |\n| `epochs` | 1-3 | Start small, increase if needed |\n\n---\n\n## Troubleshooting\n\n**Embeddings do not show improved retrieval:**\n- Verify dataset quality: triplets should have clear positive/negative distinctions\n- Use hard negatives: negatives should share some overlap with the query but not be relevant (easy negatives do not teach the model much)\n- Increase dataset size: 10K+ triplets recommended for meaningful improvement\n- Try more epochs: embedding models often need multiple passes\n- Lower learning rate: embedding models are sensitive to LR\n\n**Training loss not decreasing:**\n- Check triplet format: ensure `neg_doc` is a list even for single negatives\n- Verify hard negative quality: negatives should be challenging but clearly non-relevant\n- Increase batch size: contrastive learning benefits from larger batches\n\n**Deployment fails:**\n- Ensure you use the correct NIM image for embedding models\n- Verify sufficient GPU memory for the model size\n- Check deployment status: `client.inference.deployments.retrieve(name=deployment.name, workspace=\"default\")` and refer to platform logs for debugging\n\n## Next Steps\n\n- [Monitor training metrics](../manage-customization-jobs/get-job-status.md) in detail\n- [Evaluate your model](../../evaluator/index.md) with retrieval benchmarks\n- Integrate the fine-tuned embedding model into your RAG pipeline\n- Scale up training with the full SPECTER dataset (~684K triplets) for better results",
- "source_html": "Manual Evaluation (Recommended)
\nWhat to look for:
\nBenchmark Evaluation
\nFor systematic evaluation, use the NeMo Evaluator service with retrieval benchmarks like SciDocs, BEIR, or MTEB. Refer to the Evaluator documentation for details.
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nEmbedding-Specific Recommendations:
\n| Parameter | \nRecommended | \nNotes | \n
|---|---|---|
learning_rate | \n1e-6 to 5e-6 | \nLower than standard SFT | \n
batch_size | \n128-256 | \nLarger batches improve contrastive learning | \n
max_seq_length | \n512 | \nTypical for embedding models | \n
epochs | \n1-3 | \nStart small, increase if needed | \n
Embeddings do not show improved retrieval:
\nTraining loss not decreasing:
\nneg_doc is a list even for single negativesDeployment fails:
\nclient.inference.deployments.retrieve(name=deployment.name, workspace="default") and refer to platform logs for debuggingManual Evaluation (Recommended)
\nWhat to look for:
\nBenchmark Evaluation
\nFor systematic evaluation of end-to-end retrieval quality in a RAG pipeline, use the NeMo Evaluator RAG metrics (RAGAS context_recall, context_precision, and context_relevance).
For detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nEmbedding-Specific Recommendations:
\n| Parameter | \nRecommended | \nNotes | \n
|---|---|---|
optimizer.learning_rate | \n1e-6 to 5e-6 | \nLower than standard SFT | \n
batch.global_batch_size | \n128-256 | \nLarger batches improve contrastive learning | \n
training.max_seq_length | \n512 | \nTypical for embedding models | \n
schedule.epochs | \n1-3 | \nStart small, increase if needed | \n
Embeddings do not show improved retrieval:
\nTraining loss not decreasing:
\nneg_doc is a list even for single negativesDeployment fails:
\nclient.inference.deployments.retrieve(name=deployment.name, workspace="default") and refer to platform logs for debuggingBefore starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)The SPECTER dataset and the tutorial's base model are public and do not require a Hugging Face token. If you substitute a gated or private model, provide a token with read access.
\n" }, { "type": "markdown", @@ -69,8 +69,8 @@ export default { cells: [ }, { "type": "markdown", - "source": "### 3. Prepare Dataset\n\nUse the [SPECTER dataset](https://huggingface.co/datasets/embedding-data/SPECTER) from HuggingFace, a collection of scientific paper triplets where papers that cite each other are considered related.\n\n**Dataset structure:**\n- ~684K scientific paper triplets (this tutorial uses 10%)\n- Each triplet: (query paper, positive/related paper, negative/unrelated paper)\n- Papers that cite each other are marked as \"related\"\n\nIn this tutorial the following dataset directory structure will be used:\n```\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n```", - "source_html": "Use the SPECTER dataset from HuggingFace, a collection of scientific paper triplets where papers that cite each other are considered related.
\nDataset structure:
\nIn this tutorial the following dataset directory structure will be used:
\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n\n"
+ "source": "### 3. Prepare Dataset\n\nUse the [SPECTER dataset](https://huggingface.co/datasets/embedding-data/SPECTER) from Hugging Face, a collection of scientific paper triplets where papers that cite each other are considered related.\n\n**Dataset structure:**\n- ~684K scientific paper triplets (this tutorial uses 10%)\n- Each triplet: (query paper, positive/related paper, negative/unrelated paper)\n- Papers that cite each other are marked as \"related\"\n\nIn this tutorial the following dataset directory structure will be used:\n```\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n```",
+ "source_html": "Use the SPECTER dataset from Hugging Face, a collection of scientific paper triplets where papers that cite each other are considered related.
\nDataset structure:
\nIn this tutorial the following dataset directory structure will be used:
\nembedding-dataset\n`-- training.jsonl\n`-- validation.jsonl\n\n"
},
{
"type": "markdown",
@@ -79,9 +79,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# HuggingFace token for dataset access\nHF_TOKEN = os.environ.get(\"HF_TOKEN\")\nif not HF_TOKEN:\n raise ValueError(\"HF_TOKEN environment variable is required. Get one at https://huggingface.co/settings/tokens\")\nos.environ[\"HF_TOKEN\"] = HF_TOKEN\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path(\"embedding-dataset\").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint(\"Downloading SPECTER dataset...\")\ndata = load_dataset(\"embedding-data/SPECTER\")[\"train\"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint(\"Splitting into train/validation...\")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits[\"train\"]\nvalidation_data = splits[\"test\"]\n\n# Convert to triplet JSONL format\nprint(\"Saving to JSONL...\")\nfor name, dataset in [(\"training\", train_data), (\"validation\", validation_data)]:\n with open(f\"{DATASET_PATH}/{name}.jsonl\", \"w\") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n \"query\": row[\"set\"][0],\n \"pos_doc\": row[\"set\"][1],\n \"neg_doc\": [row[\"set\"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + \"\\n\")\n\nprint(f\"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples\")\nprint(f\"\\nExample triplet:\")\nprint(f\" Query: {train_data[0]['set'][0][:100]}...\")\nprint(f\" Positive: {train_data[0]['set'][1][:100]}...\")\nprint(f\" Negative: {train_data[0]['set'][2][:100]}...\")",
+ "source": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path(\"embedding-dataset\").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint(\"Downloading SPECTER dataset...\")\ndata = load_dataset(\"embedding-data/SPECTER\")[\"train\"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint(\"Splitting into train/validation...\")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits[\"train\"]\nvalidation_data = splits[\"test\"]\n\n# Convert to triplet JSONL format\nprint(\"Saving to JSONL...\")\nfor name, dataset in [(\"training\", train_data), (\"validation\", validation_data)]:\n with open(f\"{DATASET_PATH}/{name}.jsonl\", \"w\") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n \"query\": row[\"set\"][0],\n \"pos_doc\": row[\"set\"][1],\n \"neg_doc\": [row[\"set\"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + \"\\n\")\n\nprint(f\"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples\")\nprint(f\"\\nExample triplet:\")\nprint(f\" Query: {train_data[0]['set'][0][:100]}...\")\nprint(f\" Positive: {train_data[0]['set'][1][:100]}...\")\nprint(f\" Negative: {train_data[0]['set'][2][:100]}...\")",
"language": "python",
- "source_html": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# HuggingFace token for dataset access\nHF_TOKEN = os.environ.get("HF_TOKEN")\nif not HF_TOKEN:\n raise ValueError("HF_TOKEN environment variable is required. Get one at https://huggingface.co/settings/tokens")\nos.environ["HF_TOKEN"] = HF_TOKEN\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path("embedding-dataset").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint("Downloading SPECTER dataset...")\ndata = load_dataset("embedding-data/SPECTER")["train"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint("Splitting into train/validation...")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits["train"]\nvalidation_data = splits["test"]\n\n# Convert to triplet JSONL format\nprint("Saving to JSONL...")\nfor name, dataset in [("training", train_data), ("validation", validation_data)]:\n with open(f"{DATASET_PATH}/{name}.jsonl", "w") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n "query": row["set"][0],\n "pos_doc": row["set"][1],\n "neg_doc": [row["set"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + "\\n")\n\nprint(f"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples")\nprint(f"\\nExample triplet:")\nprint(f" Query: {train_data[0]['set'][0][:100]}...")\nprint(f" Positive: {train_data[0]['set'][1][:100]}...")\nprint(f" Negative: {train_data[0]['set'][2][:100]}...")\n"
+ "source_html": "from pathlib import Path\nfrom datasets import load_dataset\nimport json\n\n# Configuration\nDATASET_SIZE = 3000 # Number of triplets (increase for better results, max ~684K)\nVALIDATION_SPLIT = 0.05 # 5% held out for validation\nSEED = 42\nDATASET_PATH = Path("embedding-dataset").absolute()\n\n# Create directory\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Download SPECTER dataset\nprint("Downloading SPECTER dataset...")\ndata = load_dataset("embedding-data/SPECTER")["train"].shuffle(seed=SEED).select(range(DATASET_SIZE))\n\n# Split into train/validation\nprint("Splitting into train/validation...")\nsplits = data.train_test_split(test_size=VALIDATION_SPLIT, seed=SEED)\ntrain_data = splits["train"]\nvalidation_data = splits["test"]\n\n# Convert to triplet JSONL format\nprint("Saving to JSONL...")\nfor name, dataset in [("training", train_data), ("validation", validation_data)]:\n with open(f"{DATASET_PATH}/{name}.jsonl", "w") as f:\n for row in dataset:\n # SPECTER format: row['set'] = [query, positive, negative]\n triplet = {\n "query": row["set"][0],\n "pos_doc": row["set"][1],\n "neg_doc": [row["set"][2]] # List of negative documents\n }\n f.write(json.dumps(triplet) + "\\n")\n\nprint(f"\\nPrepared {len(train_data):,} training, {len(validation_data):,} validation samples")\nprint(f"\\nExample triplet:")\nprint(f" Query: {train_data[0]['set'][0][:100]}...")\nprint(f" Positive: {train_data[0]['set'][1][:100]}...")\nprint(f" Negative: {train_data[0]['set'][2][:100]}...")\n"
},
{
"type": "markdown",
@@ -96,30 +96,30 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "### 6. Secrets Setup\n\nConfigure authentication for accessing base models:\n\n- **NGC models** (`ngc://` URIs): Requires NGC API key\n- **HuggingFace models** (`hf://` URIs): Requires HF token for gated/private models\n\nGet your credentials:\n- [NGC API Key](https://ngc.nvidia.com/) (Setup → Generate API Key)\n- [HuggingFace Token](https://huggingface.co/settings/tokens) (Create token with Read access)\n\n---\n\n#### Quick Setup Example\n\nThis tutorial fine-tunes [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2), an NVIDIA embedding model optimized for question-answering and retrieval tasks.",
- "source_html": "Configure authentication for accessing base models:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsGet your credentials:
\nThis tutorial fine-tunes nvidia/llama-nemotron-embed-1b-v2, an NVIDIA embedding model optimized for question-answering and retrieval tasks.
\n" + "source": "### 6. Secrets Setup\n\nConfigure authentication for accessing base models:\n\n- **NGC models** (`ngc://` URIs): Requires NGC API key\n- **Hugging Face models** (`hf://` URIs): Requires HF token for gated/private models\n\nGet your credentials:\n- [NGC API Key](https://ngc.nvidia.com/) (Setup → Generate API Key)\n- [Hugging Face Token](https://huggingface.co/settings/tokens) (Optional; needed only for a gated/private replacement model)\n\n---\n\n#### Quick Setup Example\n\nThis tutorial fine-tunes [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2), an NVIDIA embedding model optimized for question-answering and retrieval tasks.", + "source_html": "Configure authentication for accessing base models:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsGet your credentials:
\nThis tutorial fine-tunes nvidia/llama-nemotron-embed-1b-v2, an NVIDIA embedding model optimized for question-answering and retrieval tasks.
\n" }, { "type": "code", - "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif \"NGC_SECRET_NAME\" not in globals():\n NGC_SECRET_NAME = \"ngc-api-key\"\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME}\")", + "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Public Hugging Face models need no token. Create a secret only when HF_TOKEN is set.\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\") if HF_TOKEN else None\nif hf_secret:\n print(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif \"NGC_SECRET_NAME\" not in globals():\n NGC_SECRET_NAME = \"ngc-api-key\"\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME}\")", "language": "python", - "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif "NGC_SECRET_NAME" not in globals():\n NGC_SECRET_NAME = "ngc-api-key"\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME}")\n" + "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Public Hugging Face models need no token. Create a secret only when HF_TOKEN is set.\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN") if HF_TOKEN else None\nif hf_secret:\n print(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif "NGC_SECRET_NAME" not in globals():\n NGC_SECRET_NAME = "ngc-api-key"\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME}")\n" }, { "type": "markdown", - "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2) embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.", - "source_html": "Create a fileset pointing to the nvidia/llama-nemotron-embed-1b-v2 embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n" + "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-nemotron-embed-1b-v2](https://huggingface.co/nvidia/llama-nemotron-embed-1b-v2) embedding model from Hugging Face, then create a Model Entity that references this fileset. Model downloading will take place at training time.", + "source_html": "Create a fileset pointing to the nvidia/llama-nemotron-embed-1b-v2 embedding model from Hugging Face, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n" }, { "type": "code", - "source": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"nvidia/llama-nemotron-embed-1b-v2\"\nMODEL_NAME = \"nv-nemotron-embed-1b-base\"\n\n# Ensure you have a HuggingFace token secret created\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"NVIDIA Llama Nemotron Embed 1B v2 embedding model\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n # repo_id is the full model name from Hugging Face\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n # we use the secret created in the previous step\n token_secret=hf_secret.name\n )\n )\nexcept ConflictError as e:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n\nprint(f\"\\nBase model fileset: fileset://default/{base_model.name}\")\nprint(\"\\nBase model files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=\"default\").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec}\")", + "source": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"nvidia/llama-nemotron-embed-1b-v2\"\nMODEL_NAME = \"nv-nemotron-embed-1b-base\"\n\nstorage_kwargs = {\n \"type\": \"huggingface\",\n \"repo_id\": HF_REPO_ID,\n \"repo_type\": \"model\",\n}\nif hf_secret:\n storage_kwargs[\"token_secret\"] = hf_secret.name\nstorage = HuggingfaceStorageConfigParam(**storage_kwargs)\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"NVIDIA Llama Nemotron Embed 1B v2 embedding model\",\n storage=storage,\n )\nexcept ConflictError as e:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=True,\n )\n\nprint(f\"\\nBase model fileset: fileset://default/{base_model.name}\")\nprint(\"\\nBase model files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=\"default\").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec}\")", "language": "python", - "source_html": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "nvidia/llama-nemotron-embed-1b-v2"\nMODEL_NAME = "nv-nemotron-embed-1b-base"\n\n# Ensure you have a HuggingFace token secret created\ntry:\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="NVIDIA Llama Nemotron Embed 1B v2 embedding model",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n # repo_id is the full model name from Hugging Face\n repo_id=HF_REPO_ID,\n repo_type="model",\n # we use the secret created in the previous step\n token_secret=hf_secret.name\n )\n )\nexcept ConflictError as e:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n\nprint(f"\\nBase model fileset: fileset://default/{base_model.name}")\nprint("\\nBase model files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace="default").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec}")\n" + "source_html": "import time\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "nvidia/llama-nemotron-embed-1b-v2"\nMODEL_NAME = "nv-nemotron-embed-1b-base"\n\nstorage_kwargs = {\n "type": "huggingface",\n "repo_id": HF_REPO_ID,\n "repo_type": "model",\n}\nif hf_secret:\n storage_kwargs["token_secret"] = hf_secret.name\nstorage = HuggingfaceStorageConfigParam(**storage_kwargs)\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="NVIDIA Llama Nemotron Embed 1B v2 embedding model",\n storage=storage,\n )\nexcept ConflictError as e:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=True,\n )\n\nprint(f"\\nBase model fileset: fileset://default/{base_model.name}")\nprint("\\nBase model files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace="default").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec}")\n" }, { "type": "markdown", - "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with split `schedule`, `batch`, `optimizer`, and `parallelism` sections. Reference the model entity and dataset fileset by workspace/name (not `fileset://` URIs).\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training.training_type`**: `sft`\n- **`training.finetuning_type`**: `all_weights` for full fine-tuning, or `lora_merged` for merged LoRA\n- **`optimizer.learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch.global_batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set `finetuning_type` to `lora_merged`:\n\n```python\ntraining={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora_merged\",\n \"lora\": {\"rank\": 16, \"alpha\": 32},\n \"max_seq_length\": MAX_SEQ_LENGTH,\n}\n```", - "source_html": "Create a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nSubmit to the Automodel backend using AutomodelJobInput with split schedule, batch, optimizer, and parallelism sections. Reference the model entity and dataset fileset by workspace/name (not fileset:// URIs).
Key hyperparameters for embedding fine-tuning:
\ntraining.training_type: sfttraining.finetuning_type: all_weights for full fine-tuning, or lora_merged for merged LoRAoptimizer.learning_rate: Lower values (1e-6 to 5e-6) work well for embedding modelsbatch.global_batch_size: Larger batches improve contrastive learning (128-256 recommended)NOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set finetuning_type to lora_merged:
training={\n "training_type": "sft",\n "finetuning_type": "lora_merged",\n "lora": {"rank": 16, "alpha": 32},\n "max_seq_length": MAX_SEQ_LENGTH,\n}\n\n"
+ "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with split `schedule`, `batch`, `optimizer`, and `parallelism` sections. Reference the model entity and dataset fileset by workspace/name (not `fileset://` URIs).\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training.training_type`**: `sft`\n- **`training.finetuning_type`**: `all_weights` for full fine-tuning, or `lora_merged` for merged LoRA\n- **`optimizer.learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch.global_batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set `finetuning_type` to `lora_merged`:\n\n```python\ntraining={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora_merged\",\n \"lora\": {\"rank\": 16, \"alpha\": 32},\n \"max_seq_length\": 512,\n}\n```",
+ "source_html": "Create a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nSubmit to the Automodel backend using AutomodelJobInput with split schedule, batch, optimizer, and parallelism sections. Reference the model entity and dataset fileset by workspace/name (not fileset:// URIs).
Key hyperparameters for embedding fine-tuning:
\ntraining.training_type: sfttraining.finetuning_type: all_weights for full fine-tuning, or lora_merged for merged LoRAoptimizer.learning_rate: Lower values (1e-6 to 5e-6) work well for embedding modelsbatch.global_batch_size: Larger batches improve contrastive learning (128-256 recommended)NOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set finetuning_type to lora_merged:
training={\n "training_type": "sft",\n "finetuning_type": "lora_merged",\n "lora": {"rank": 16, "alpha": 32},\n "max_seq_length": 512,\n}\n\n"
},
{
"type": "code",
@@ -134,9 +134,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
+ "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)\n\nif status.status != \"completed\":\n raise RuntimeError(f\"Training job finished with status: {status.status}\")",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n\nif status.status != "completed":\n raise RuntimeError(f"Training job finished with status: {status.status}")\n"
},
{
"type": "markdown",
@@ -184,7 +184,7 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world queries from your domain\n- Compare retrieval rankings before and after fine-tuning\n- Check that semantically similar items rank higher than keyword matches\n\n**What to look for:**\n- ✅ Relevant documents consistently rank in top positions\n- ✅ Keyword traps (like \"Random Forest\" vs \"Random Fields\") are handled correctly\n- ✅ Domain-specific terminology is understood\n- ❌ Unrelated documents with matching keywords do not rank high\n\n**Benchmark Evaluation**\n\nFor systematic evaluation, use the NeMo Evaluator service with retrieval benchmarks like SciDocs, BEIR, or MTEB. Refer to the [Evaluator documentation](../../evaluator/index.md) for details.\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n**Embedding-Specific Recommendations:**\n\n| Parameter | Recommended | Notes |\n|-----------|-------------|-------|\n| `learning_rate` | 1e-6 to 5e-6 | Lower than standard SFT |\n| `batch_size` | 128-256 | Larger batches improve contrastive learning |\n| `max_seq_length` | 512 | Typical for embedding models |\n| `epochs` | 1-3 | Start small, increase if needed |\n\n---\n\n## Troubleshooting\n\n**Embeddings do not show improved retrieval:**\n- Verify dataset quality: triplets should have clear positive/negative distinctions\n- Use hard negatives: negatives should share some overlap with the query but not be relevant (easy negatives do not teach the model much)\n- Increase dataset size: 10K+ triplets recommended for meaningful improvement\n- Try more epochs: embedding models often need multiple passes\n- Lower learning rate: embedding models are sensitive to LR\n\n**Training loss not decreasing:**\n- Check triplet format: ensure `neg_doc` is a list even for single negatives\n- Verify hard negative quality: negatives should be challenging but clearly non-relevant\n- Increase batch size: contrastive learning benefits from larger batches\n\n**Deployment fails:**\n- Ensure you use the correct NIM image for embedding models\n- Verify sufficient GPU memory for the model size\n- Check deployment status: `client.inference.deployments.retrieve(name=deployment.name, workspace=\"default\")` and refer to platform logs for debugging\n\n## Next Steps\n\n- [Monitor training metrics](../manage-customization-jobs/get-job-status.md) in detail\n- [Evaluate your model](../../evaluator/index.md) with retrieval benchmarks\n- Integrate the fine-tuned embedding model into your RAG pipeline\n- Scale up training with the full SPECTER dataset (~684K triplets) for better results",
- "source_html": "Manual Evaluation (Recommended)
\nWhat to look for:
\nBenchmark Evaluation
\nFor systematic evaluation, use the NeMo Evaluator service with retrieval benchmarks like SciDocs, BEIR, or MTEB. Refer to the Evaluator documentation for details.
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nEmbedding-Specific Recommendations:
\n| Parameter | \nRecommended | \nNotes | \n
|---|---|---|
learning_rate | \n1e-6 to 5e-6 | \nLower than standard SFT | \n
batch_size | \n128-256 | \nLarger batches improve contrastive learning | \n
max_seq_length | \n512 | \nTypical for embedding models | \n
epochs | \n1-3 | \nStart small, increase if needed | \n
Embeddings do not show improved retrieval:
\nTraining loss not decreasing:
\nneg_doc is a list even for single negativesDeployment fails:
\nclient.inference.deployments.retrieve(name=deployment.name, workspace="default") and refer to platform logs for debuggingManual Evaluation (Recommended)
\nWhat to look for:
\nBenchmark Evaluation
\nFor systematic evaluation of end-to-end retrieval quality in a RAG pipeline, use the NeMo Evaluator RAG metrics (RAGAS context_recall, context_precision, and context_relevance).
For detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nEmbedding-Specific Recommendations:
\n| Parameter | \nRecommended | \nNotes | \n
|---|---|---|
optimizer.learning_rate | \n1e-6 to 5e-6 | \nLower than standard SFT | \n
batch.global_batch_size | \n128-256 | \nLarger batches improve contrastive learning | \n
training.max_seq_length | \n512 | \nTypical for embedding models | \n
schedule.epochs | \n1-3 | \nStart small, increase if needed | \n
Embeddings do not show improved retrieval:
\nTraining loss not decreasing:
\nneg_doc is a list even for single negativesDeployment fails:
\nclient.inference.deployments.retrieve(name=deployment.name, workspace="default") and refer to platform logs for debuggingBefore starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)datasets package for loading SQuAD: pip install datasetsBefore starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)datasets package for loading SQuAD: pip install datasetsFor Huggingface models that require authentication, create a secret with your HF token. Get a token from Huggingface Settings and accept the model terms.
\nThis is generally true for LLaMa based models (e.g. Llama-3.2-1B-Instruct).
\nexport HF_TOKEN=<your-huggingface-token>\n\n"
+ "source": "### 4. Secrets Setup\n\nFor Hugging Face models that require authentication, create a secret with your HF token. Get a token from [Hugging Face Settings](https://huggingface.co/settings/tokens) and accept the model terms.\n\nThis is generally true for Llama-based models (for example, [Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)).\n\n```sh\nexport HF_TOKEN=For Hugging Face models that require authentication, create a secret with your HF token. Get a token from Hugging Face Settings and accept the model terms.
\nThis is generally true for Llama-based models (for example, Llama-3.2-1B-Instruct).
\nexport HF_TOKEN=<your-huggingface-token>\n\n"
},
{
"type": "code",
@@ -71,9 +71,9 @@
},
{
"type": "code",
- "source": "HF_REPO_ID = \"Qwen/Qwen3-0.6B\"\nMODEL_NAME = \"qwen3-0.6b\"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n )\n if hf_secret:\n storage[\"token_secret\"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"Qwen3 0.6b base model from Huggingface\",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace=\"default\", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\nprint(f\"Base model fileset: fileset://default/{base_model.name}\")\nprint(client.files.list(fileset=MODEL_NAME, workspace=\"default\"))\n\ntime_check = max_wait_time_checker(600, \"Model Spec\")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f\"ModelSpec: {base_model.spec}\")",
+ "source": "HF_REPO_ID = \"Qwen/Qwen3-0.6B\"\nMODEL_NAME = \"qwen3-0.6b\"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n )\n if hf_secret:\n storage[\"token_secret\"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"Qwen3 0.6b base model from Hugging Face\",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace=\"default\", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\nprint(f\"Base model fileset: fileset://default/{base_model.name}\")\nprint(client.files.list(fileset=MODEL_NAME, workspace=\"default\"))\n\ntime_check = max_wait_time_checker(600, \"Model Spec\")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f\"ModelSpec: {base_model.spec}\")",
"language": "python",
- "source_html": "HF_REPO_ID = "Qwen/Qwen3-0.6B"\nMODEL_NAME = "qwen3-0.6b"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n )\n if hf_secret:\n storage["token_secret"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="Qwen3 0.6b base model from Huggingface",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace="default", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\nprint(f"Base model fileset: fileset://default/{base_model.name}")\nprint(client.files.list(fileset=MODEL_NAME, workspace="default"))\n\ntime_check = max_wait_time_checker(600, "Model Spec")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f"ModelSpec: {base_model.spec}")\n"
+ "source_html": "HF_REPO_ID = "Qwen/Qwen3-0.6B"\nMODEL_NAME = "qwen3-0.6b"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n )\n if hf_secret:\n storage["token_secret"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="Qwen3 0.6b base model from Hugging Face",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace="default", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\nprint(f"Base model fileset: fileset://default/{base_model.name}")\nprint(client.files.list(fileset=MODEL_NAME, workspace="default"))\n\ntime_check = max_wait_time_checker(600, "Model Spec")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f"ModelSpec: {base_model.spec}")\n"
},
{
"type": "markdown",
@@ -126,9 +126,9 @@
},
{
"type": "code",
- "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": OUTPUT_NAME,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
+ "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nINFERENCE_MODEL_NAME = f\"default--{OUTPUT_NAME}\"\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": INFERENCE_MODEL_NAME,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
"language": "python",
- "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": OUTPUT_NAME,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
+ "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nINFERENCE_MODEL_NAME = f"default--{OUTPUT_NAME}"\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": INFERENCE_MODEL_NAME,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/lora-customization-job.ts b/docs/fern/components/notebooks/lora-customization-job.ts
index d4b0272cd4..dd7691c96a 100644
--- a/docs/fern/components/notebooks/lora-customization-job.ts
+++ b/docs/fern/components/notebooks/lora-customization-job.ts
@@ -12,8 +12,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **Installed the `datasets` package** for loading SQuAD: `pip install datasets`\n4. **At least one GPU with CUDA 12.8+**",
- "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)datasets package for loading SQuAD: pip install datasetsBefore starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)datasets package for loading SQuAD: pip install datasetsFor Huggingface models that require authentication, create a secret with your HF token. Get a token from Huggingface Settings and accept the model terms.
\nThis is generally true for LLaMa based models (e.g. Llama-3.2-1B-Instruct).
\nexport HF_TOKEN=<your-huggingface-token>\n\n"
+ "source": "### 4. Secrets Setup\n\nFor Hugging Face models that require authentication, create a secret with your HF token. Get a token from [Hugging Face Settings](https://huggingface.co/settings/tokens) and accept the model terms.\n\nThis is generally true for Llama-based models (for example, [Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)).\n\n```sh\nexport HF_TOKEN=For Hugging Face models that require authentication, create a secret with your HF token. Get a token from Hugging Face Settings and accept the model terms.
\nThis is generally true for Llama-based models (for example, Llama-3.2-1B-Instruct).
\nexport HF_TOKEN=<your-huggingface-token>\n\n"
},
{
"type": "code",
@@ -76,9 +76,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "HF_REPO_ID = \"Qwen/Qwen3-0.6B\"\nMODEL_NAME = \"qwen3-0.6b\"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n )\n if hf_secret:\n storage[\"token_secret\"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"Qwen3 0.6b base model from Huggingface\",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace=\"default\", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\nprint(f\"Base model fileset: fileset://default/{base_model.name}\")\nprint(client.files.list(fileset=MODEL_NAME, workspace=\"default\"))\n\ntime_check = max_wait_time_checker(600, \"Model Spec\")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f\"ModelSpec: {base_model.spec}\")",
+ "source": "HF_REPO_ID = \"Qwen/Qwen3-0.6B\"\nMODEL_NAME = \"qwen3-0.6b\"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n )\n if hf_secret:\n storage[\"token_secret\"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"Qwen3 0.6b base model from Hugging Face\",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace=\"default\", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\nprint(f\"Base model fileset: fileset://default/{base_model.name}\")\nprint(client.files.list(fileset=MODEL_NAME, workspace=\"default\"))\n\ntime_check = max_wait_time_checker(600, \"Model Spec\")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f\"ModelSpec: {base_model.spec}\")",
"language": "python",
- "source_html": "HF_REPO_ID = "Qwen/Qwen3-0.6B"\nMODEL_NAME = "qwen3-0.6b"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n )\n if hf_secret:\n storage["token_secret"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="Qwen3 0.6b base model from Huggingface",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace="default", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\nprint(f"Base model fileset: fileset://default/{base_model.name}")\nprint(client.files.list(fileset=MODEL_NAME, workspace="default"))\n\ntime_check = max_wait_time_checker(600, "Model Spec")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f"ModelSpec: {base_model.spec}")\n"
+ "source_html": "HF_REPO_ID = "Qwen/Qwen3-0.6B"\nMODEL_NAME = "qwen3-0.6b"\n\ntry:\n storage = HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n )\n if hf_secret:\n storage["token_secret"] = hf_secret.name\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="Qwen3 0.6b base model from Hugging Face",\n storage=storage,\n cache=True,\n )\nexcept ConflictError:\n base_model_fs = client.files.filesets.retrieve(workspace="default", name=MODEL_NAME)\n\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\nexcept ConflictError:\n client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n trust_remote_code=False,\n )\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\nprint(f"Base model fileset: fileset://default/{base_model.name}")\nprint(client.files.list(fileset=MODEL_NAME, workspace="default"))\n\ntime_check = max_wait_time_checker(600, "Model Spec")\nwhile not base_model.spec:\n time_check()\n time.sleep(10)\n base_model = client.models.retrieve(workspace="default", name=MODEL_NAME)\n\n# Clear verbose linear_layers list for cleaner output\nbase_model.spec.linear_layers = None\nprint(f"ModelSpec: {base_model.spec}")\n"
},
{
"type": "markdown",
@@ -131,9 +131,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": OUTPUT_NAME,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
+ "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nINFERENCE_MODEL_NAME = f\"default--{OUTPUT_NAME}\"\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": INFERENCE_MODEL_NAME,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
"language": "python",
- "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": OUTPUT_NAME,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
+ "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nINFERENCE_MODEL_NAME = f"default--{OUTPUT_NAME}"\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": INFERENCE_MODEL_NAME,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/optimize-throughput.json b/docs/fern/components/notebooks/optimize-throughput.json
index caa20e705d..c4d4b90d72 100644
--- a/docs/fern/components/notebooks/optimize-throughput.json
+++ b/docs/fern/components/notebooks/optimize-throughput.json
@@ -7,8 +7,8 @@
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)",
- "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)If you plan to use NGC or HuggingFace models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nThis tutorial uses the meta-llama/Llama-3.2-1B-Instruct model from HuggingFace. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access.
\nHuggingFace Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for the model in the next step.If you plan to use NGC or Hugging Face models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nThis tutorial uses the meta-llama/Llama-3.2-1B-Instruct model from Hugging Face. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access.
\nHugging Face Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for the model in the next step.Create a fileset pointing to the meta-llama/Llama-3.2-1B-Instruct model on HuggingFace. This step creates a pointer to the model on Hugging Face and does not download it. The model is downloaded at job creation time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
Create a fileset pointing to the meta-llama/Llama-3.2-1B-Instruct model on Hugging Face. This step creates a pointer to the model on Hugging Face and does not download it. The model is downloaded at job creation time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
A training job contains multiple steps:
\nThe elapsed time printed below reflects progress of the entire job. We compare the time taken by the finetuning step for both jobs in the last section of this tutorial.
\n" + "source": "### 6. Track Fine-Tuning Progress\n\nA training job contains multiple steps: \n- Model and dataset downloading\n- Fine-tuning where LoRA adapter weights are trained\n- Creating a fileset entry for the fine-tuned model\n- Fine-tuned weights uploading\n\nThe elapsed time printed below reflects progress of the entire job. We compare the time taken by the fine-tuning step for both jobs in the last section of this tutorial.", + "source_html": "A training job contains multiple steps:
\nThe elapsed time printed below reflects progress of the entire job. We compare the time taken by the fine-tuning step for both jobs in the last section of this tutorial.
\n" }, { "type": "markdown", @@ -105,14 +105,14 @@ }, { "type": "markdown", - "source": "#### Monitor the Job Until Completion\n\nThe cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the finetuning step begins.\n\n> **Note:** This is additional code. You can also use the Weights & Biases or MLflow integrations.", - "source_html": "The cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the finetuning step begins.
\n\n\n" + "source": "#### Monitor the Job Until Completion\n\nThe cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the fine-tuning step begins.\n\n> **Note:** This is additional code. You can also use the Weights & Biases or MLflow integrations.", + "source_html": "Note: This is additional code. You can also use the Weights & Biases or MLflow integrations.
\n
The cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the fine-tuning step begins.
\n\n\n" }, { "type": "code", - "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"loss\"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get(\"val_loss\")\nif packed_val_loss is not None:\n print(f\"Validation loss: {float(packed_val_loss):.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")", + "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"train_loss\"\n\n\ndef get_training_metric(\n status: PlatformJobStatusResponse,\n metric_key: str,\n) -> float | None:\n \"\"\"Return a metric reported by a task in the training step.\"\"\"\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n value = (task.status_details or {}).get(metric_key)\n if value is not None:\n return float(value)\n return None\n\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = get_training_metric(job_with_sequence_packing_status, VAL_LOSS_KEY)\nif packed_val_loss is not None:\n print(f\"Validation loss: {packed_val_loss:.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")", "language": "python", - "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "loss"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "training":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get("val_loss")\nif packed_val_loss is not None:\n print(f"Validation loss: {float(packed_val_loss):.2f}")\nelse:\n print("Validation loss: not reported in job status")\n" + "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "train_loss"\n\n\ndef get_training_metric(\n status: PlatformJobStatusResponse,\n metric_key: str,\n) -> float | None:\n """Return a metric reported by a task in the training step."""\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n value = (task.status_details or {}).get(metric_key)\n if value is not None:\n return float(value)\n return None\n\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "training":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = get_training_metric(job_with_sequence_packing_status, VAL_LOSS_KEY)\nif packed_val_loss is not None:\n print(f"Validation loss: {packed_val_loss:.2f}")\nelse:\n print("Validation loss: not reported in job status")\n" }, { "type": "markdown", @@ -127,14 +127,14 @@ }, { "type": "markdown", - "source": "### 8. Track Finetuning Progress for Job without Sequence Packing", - "source_html": "Note: This is additional code. You can also use the Weights & Biases or MLflow integrations.
\n
The expected validation loss curves should match closely for both jobs.\n
Sequence packed version should complete significantly faster.\n
Sequence packed version should have a higher GPU utilization.\n
Sequence packed version should have a higher GPU Memory Allocation.\n
The expected validation loss curves should match closely for both jobs.\n
Sequence packed version should complete significantly faster.\n
Sequence packed version should have a higher GPU utilization.\n
Sequence packed version should have a higher GPU Memory Allocation.\n
Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)If you plan to use NGC or HuggingFace models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nThis tutorial uses the meta-llama/Llama-3.2-1B-Instruct model from HuggingFace. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access.
\nHuggingFace Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for the model in the next step.If you plan to use NGC or Hugging Face models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nThis tutorial uses the meta-llama/Llama-3.2-1B-Instruct model from Hugging Face. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access.
\nHugging Face Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for the model in the next step.Create a fileset pointing to the meta-llama/Llama-3.2-1B-Instruct model on HuggingFace. This step creates a pointer to the model on Hugging Face and does not download it. The model is downloaded at job creation time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
Create a fileset pointing to the meta-llama/Llama-3.2-1B-Instruct model on Hugging Face. This step creates a pointer to the model on Hugging Face and does not download it. The model is downloaded at job creation time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
A training job contains multiple steps:
\nThe elapsed time printed below reflects progress of the entire job. We compare the time taken by the finetuning step for both jobs in the last section of this tutorial.
\n" + "source": "### 6. Track Fine-Tuning Progress\n\nA training job contains multiple steps: \n- Model and dataset downloading\n- Fine-tuning where LoRA adapter weights are trained\n- Creating a fileset entry for the fine-tuned model\n- Fine-tuned weights uploading\n\nThe elapsed time printed below reflects progress of the entire job. We compare the time taken by the fine-tuning step for both jobs in the last section of this tutorial.", + "source_html": "A training job contains multiple steps:
\nThe elapsed time printed below reflects progress of the entire job. We compare the time taken by the fine-tuning step for both jobs in the last section of this tutorial.
\n" }, { "type": "markdown", @@ -110,14 +110,14 @@ export default { cells: [ }, { "type": "markdown", - "source": "#### Monitor the Job Until Completion\n\nThe cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the finetuning step begins.\n\n> **Note:** This is additional code. You can also use the Weights & Biases or MLflow integrations.", - "source_html": "The cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the finetuning step begins.
\n\n\n" + "source": "#### Monitor the Job Until Completion\n\nThe cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the fine-tuning step begins.\n\n> **Note:** This is additional code. You can also use the Weights & Biases or MLflow integrations.", + "source_html": "Note: This is additional code. You can also use the Weights & Biases or MLflow integrations.
\n
The cell below polls the job status every 10 seconds and renders a live dashboard with validation loss, GPU VRAM usage, and GPU utilization charts. The charts appear empty at first while the model and dataset download; training metrics and GPU activity populate after the fine-tuning step begins.
\n\n\n" }, { "type": "code", - "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"loss\"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get(\"val_loss\")\nif packed_val_loss is not None:\n print(f\"Validation loss: {float(packed_val_loss):.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")", + "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"train_loss\"\n\n\ndef get_training_metric(\n status: PlatformJobStatusResponse,\n metric_key: str,\n) -> float | None:\n \"\"\"Return a metric reported by a task in the training step.\"\"\"\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n value = (task.status_details or {}).get(metric_key)\n if value is not None:\n return float(value)\n return None\n\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = get_training_metric(job_with_sequence_packing_status, VAL_LOSS_KEY)\nif packed_val_loss is not None:\n print(f\"Validation loss: {packed_val_loss:.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")", "language": "python", - "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "loss"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "training":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get("val_loss")\nif packed_val_loss is not None:\n print(f"Validation loss: {float(packed_val_loss):.2f}")\nelse:\n print("Validation loss: not reported in job status")\n" + "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "train_loss"\n\n\ndef get_training_metric(\n status: PlatformJobStatusResponse,\n metric_key: str,\n) -> float | None:\n """Return a metric reported by a task in the training step."""\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n value = (task.status_details or {}).get(metric_key)\n if value is not None:\n return float(value)\n return None\n\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "training":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = get_training_metric(job_with_sequence_packing_status, VAL_LOSS_KEY)\nif packed_val_loss is not None:\n print(f"Validation loss: {packed_val_loss:.2f}")\nelse:\n print("Validation loss: not reported in job status")\n" }, { "type": "markdown", @@ -132,14 +132,14 @@ export default { cells: [ }, { "type": "markdown", - "source": "### 8. Track Finetuning Progress for Job without Sequence Packing", - "source_html": "Note: This is additional code. You can also use the Weights & Biases or MLflow integrations.
\n
The expected validation loss curves should match closely for both jobs.\n
Sequence packed version should complete significantly faster.\n
Sequence packed version should have a higher GPU utilization.\n
Sequence packed version should have a higher GPU Memory Allocation.\n
The expected validation loss curves should match closely for both jobs.\n
Sequence packed version should complete significantly faster.\n
Sequence packed version should have a higher GPU utilization.\n
Sequence packed version should have a higher GPU Memory Allocation.\n
Learn how to fine-tune all model weights using supervised fine-tuning (SFT) to customize LLM behavior for your specific tasks.
\nSupervised Fine-Tuning (SFT) customizes model behavior, injects new knowledge, and optimizes performance for specific domains and tasks. Full SFT modifies all model weights during training, providing maximum customization flexibility.
\nWhat you can achieve with SFT:
\nFull SFT trains all model parameters (for example, all 70 billion weights in Llama 70B):
\nLoRA trains only ~1% of weights by adding thin matrices to existing weights:
\nWhen to choose Full SFT:
\nWhen to choose LoRA: Refer to the LoRA tutorial for most use cases, especially with large models (70B+) or limited GPU resources.
\n" + "source": "\n\n\n# Full SFT Customization\n\nLearn how to fine-tune all model weights using supervised fine-tuning (SFT) to customize LLM behavior for your specific tasks.\n\n## About\n\nSupervised Fine-Tuning (SFT) customizes model behavior, injects new knowledge, and optimizes performance for specific domains and tasks. Full SFT modifies **all model weights** during training, providing maximum customization flexibility.\n\n**What you can achieve with SFT:**\n\n- 🎯 **Specialize for domains:** Fine-tune models on legal texts, medical records, or financial data\n- 💡 **Inject knowledge:** Add new information not present in the base model\n- 📈 **Improve accuracy:** Optimize for specific tasks like sentiment analysis, summarization, or code generation\n\n### SFT vs LoRA: Understanding the Trade-offs\n\n**Full SFT** trains all model parameters (for example, all 70 billion weights in Llama 70B):\n\n- ✅ Maximum model adaptation and knowledge injection\n- ✅ Can fundamentally change model behavior\n- ✅ Best for significant domain shifts or specialized tasks\n- ❌ Requires substantial GPU resources (4-8x more than LoRA)\n- ❌ Produces a full BF16 checkpoint (~140 GB for Llama 70B); peak job disk usage can reach approximately 3× the downloaded base checkpoint size\n- ❌ Longer training time\n\n**LoRA** trains only ~1% of weights by adding thin matrices to existing weights:\n\n- ✅ 75-95% less memory required\n- ✅ Faster training (2-4x speedup)\n- ✅ Produces small adapter files (~100-500MB)\n- ✅ Multiple adapters can share one base model\n- ❌ Limited adaptation capability compared to full fine-tuning\n\n**When to choose Full SFT:**\n\n- Training small models (1B-8B) where resource cost is manageable\n- Need fundamental behavior changes (for example, medical diagnosis, legal reasoning)\n- Injecting substantial new knowledge not in the base model\n\n**When to choose LoRA:** Refer to the [LoRA tutorial](./lora-customization-job) for most use cases, especially with large models (70B+) or limited GPU resources.", + "source_html": "\n\nLearn how to fine-tune all model weights using supervised fine-tuning (SFT) to customize LLM behavior for your specific tasks.
\nSupervised Fine-Tuning (SFT) customizes model behavior, injects new knowledge, and optimizes performance for specific domains and tasks. Full SFT modifies all model weights during training, providing maximum customization flexibility.
\nWhat you can achieve with SFT:
\nFull SFT trains all model parameters (for example, all 70 billion weights in Llama 70B):
\nLoRA trains only ~1% of weights by adding thin matrices to existing weights:
\nWhen to choose Full SFT:
\nWhen to choose LoRA: Refer to the LoRA tutorial for most use cases, especially with large models (70B+) or limited GPU resources.
\n" }, { "type": "markdown", - "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)", - "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)If you plan to use NGC or HuggingFace models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nIn this tutorial we are going to work with meta-llama/Llama-3.2-1B-Instruct model from HuggingFace. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access
\nHuggingFace Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for model in the next stepIf you plan to use NGC or Hugging Face models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nIn this tutorial we are going to work with the meta-llama/Llama-3.2-1B-Instruct model from Hugging Face. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access.
\nHugging Face Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for model in the next stepCreate a fileset pointing to meta-llama/Llama-3.2-1B-Instruct model in HuggingFace that we will train with SFT. Then create a Model Entity that references this fileset. Model downloading will take place at training time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
Create a fileset pointing to the meta-llama/Llama-3.2-1B-Instruct model in Hugging Face that we will train with SFT. Then create a Model Entity that references this fileset. Model downloading will take place at training time.
\nThis tutorial's default model is gated, so the fileset includes token_secret=hf_secret.name. If you substitute a public model, you can omit token_secret.
Create a customization job to fine-tune all model weights using the Automodel backend and AutomodelJobInput.
Create a customization job to fine-tune all model weights using the Automodel backend and AutomodelJobInput.
Manual Evaluation (Recommended)
\nWhat to look for:
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\ntoken_secret=hf_secret.name for gated modelsAutomodelJobInput references use the workspace/name format: model=f"default/{MODEL_NAME}" and dataset={"training": f"default/{DATASET_NAME}"} (for example, default/llama-3-2-1b-base, default/sft-dataset)fileset=f"default/{MODEL_NAME}"client.jobs.get_status(name=job.job.name, workspace="default")Job fails with OOM (Out of Memory) error:
\nglobal_batch_size from 64 to 32 or 16 in batch={...}micro_batch_size at 1 (already the minimum in this tutorial)max_seq_length from 2048 to 1024 or 512 in training={...}num_gpus_per_node and tensor_parallel_size in parallelism={...}Loss curves not decreasing (underfitting):
\nepochs from 2 to 3-5 in schedule={...}1e-4 or 1e-5 instead of the default 5e-5 in optimizer={...}Training loss decreases but validation loss increases (overfitting):
\nepochs from 2 to 1 in schedule={...}learning_rate from 5e-5 to 2e-5 or 1e-5 in optimizer={...}Model output quality is poor despite good training metrics:
\nlearning_rate and global_batch_sizeDeployment fails:
\nclient.models.retrieve(name=OUTPUT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")executor_config={"gpu": 1, ...}engine="vllm" with vllm/vllm-openai:v0.22.1Manual Evaluation (Recommended)
\nWhat to look for:
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\ntoken_secret=hf_secret.name for gated modelsAutomodelJobInput references use the workspace/name format: model=f"default/{MODEL_NAME}" and dataset={"training": f"default/{DATASET_NAME}"} (for example, default/llama-3-2-1b-base, default/sft-dataset)fileset=f"default/{MODEL_NAME}"client.jobs.get_status(name=job.job.name, workspace="default")Job fails with OOM (Out of Memory) error:
\nglobal_batch_size from 64 to 32 or 16 in batch={...}micro_batch_size at 1 (already the minimum in this tutorial)max_seq_length from 2048 to 1024 or 512 in training={...}num_gpus_per_node and tensor_parallel_size in parallelism={...}Loss curves not decreasing (underfitting):
\nepochs from 2 to 3-5 in schedule={...}1e-4 or 1e-5 instead of the default 5e-5 in optimizer={...}Training loss decreases but validation loss increases (overfitting):
\nepochs from 2 to 1 in schedule={...}learning_rate from 5e-5 to 2e-5 or 1e-5 in optimizer={...}Model output quality is poor despite good training metrics:
\nlearning_rate and global_batch_sizeDeployment fails:
\nclient.models.retrieve(name=OUTPUT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")executor_config={"gpu": 1, ...}engine="vllm" with vllm/vllm-openai:v0.22.1Learn how to fine-tune all model weights using supervised fine-tuning (SFT) to customize LLM behavior for your specific tasks.
\nSupervised Fine-Tuning (SFT) customizes model behavior, injects new knowledge, and optimizes performance for specific domains and tasks. Full SFT modifies all model weights during training, providing maximum customization flexibility.
\nWhat you can achieve with SFT:
\nFull SFT trains all model parameters (for example, all 70 billion weights in Llama 70B):
\nLoRA trains only ~1% of weights by adding thin matrices to existing weights:
\nWhen to choose Full SFT:
\nWhen to choose LoRA: Refer to the LoRA tutorial for most use cases, especially with large models (70B+) or limited GPU resources.
\n" + "source": "\n\n\n# Full SFT Customization\n\nLearn how to fine-tune all model weights using supervised fine-tuning (SFT) to customize LLM behavior for your specific tasks.\n\n## About\n\nSupervised Fine-Tuning (SFT) customizes model behavior, injects new knowledge, and optimizes performance for specific domains and tasks. Full SFT modifies **all model weights** during training, providing maximum customization flexibility.\n\n**What you can achieve with SFT:**\n\n- 🎯 **Specialize for domains:** Fine-tune models on legal texts, medical records, or financial data\n- 💡 **Inject knowledge:** Add new information not present in the base model\n- 📈 **Improve accuracy:** Optimize for specific tasks like sentiment analysis, summarization, or code generation\n\n### SFT vs LoRA: Understanding the Trade-offs\n\n**Full SFT** trains all model parameters (for example, all 70 billion weights in Llama 70B):\n\n- ✅ Maximum model adaptation and knowledge injection\n- ✅ Can fundamentally change model behavior\n- ✅ Best for significant domain shifts or specialized tasks\n- ❌ Requires substantial GPU resources (4-8x more than LoRA)\n- ❌ Produces a full BF16 checkpoint (~140 GB for Llama 70B); peak job disk usage can reach approximately 3× the downloaded base checkpoint size\n- ❌ Longer training time\n\n**LoRA** trains only ~1% of weights by adding thin matrices to existing weights:\n\n- ✅ 75-95% less memory required\n- ✅ Faster training (2-4x speedup)\n- ✅ Produces small adapter files (~100-500MB)\n- ✅ Multiple adapters can share one base model\n- ❌ Limited adaptation capability compared to full fine-tuning\n\n**When to choose Full SFT:**\n\n- Training small models (1B-8B) where resource cost is manageable\n- Need fundamental behavior changes (for example, medical diagnosis, legal reasoning)\n- Injecting substantial new knowledge not in the base model\n\n**When to choose LoRA:** Refer to the [LoRA tutorial](./lora-customization-job) for most use cases, especially with large models (70B+) or limited GPU resources.", + "source_html": "\n\nLearn how to fine-tune all model weights using supervised fine-tuning (SFT) to customize LLM behavior for your specific tasks.
\nSupervised Fine-Tuning (SFT) customizes model behavior, injects new knowledge, and optimizes performance for specific domains and tasks. Full SFT modifies all model weights during training, providing maximum customization flexibility.
\nWhat you can achieve with SFT:
\nFull SFT trains all model parameters (for example, all 70 billion weights in Llama 70B):
\nLoRA trains only ~1% of weights by adding thin matrices to existing weights:
\nWhen to choose Full SFT:
\nWhen to choose LoRA: Refer to the LoRA tutorial for most use cases, especially with large models (70B+) or limited GPU resources.
\n" }, { "type": "markdown", - "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)", - "source_html": "Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)Before starting this tutorial, ensure you have:
\npip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root)If you plan to use NGC or HuggingFace models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nIn this tutorial we are going to work with meta-llama/Llama-3.2-1B-Instruct model from HuggingFace. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access
\nHuggingFace Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for model in the next stepIf you plan to use NGC or Hugging Face models, you will need to configure authentication:
\nngc:// URIs): Requires NGC API keyhf:// URIs): Requires HF token for gated/private modelsConfigure these as secrets in your platform. Refer to Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\nIn this tutorial we are going to work with the meta-llama/Llama-3.2-1B-Instruct model from Hugging Face. Ensure that you have sufficient permissions to download the model. If you cannot access the files on the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access.
\nHugging Face Authentication:
\ntoken_secret parametertoken_secret parameter when creating a fileset for model in the next stepCreate a fileset pointing to meta-llama/Llama-3.2-1B-Instruct model in HuggingFace that we will train with SFT. Then create a Model Entity that references this fileset. Model downloading will take place at training time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
Create a fileset pointing to the meta-llama/Llama-3.2-1B-Instruct model in Hugging Face that we will train with SFT. Then create a Model Entity that references this fileset. Model downloading will take place at training time.
\nThis tutorial's default model is gated, so the fileset includes token_secret=hf_secret.name. If you substitute a public model, you can omit token_secret.
Create a customization job to fine-tune all model weights using the Automodel backend and AutomodelJobInput.
Create a customization job to fine-tune all model weights using the Automodel backend and AutomodelJobInput.
Manual Evaluation (Recommended)
\nWhat to look for:
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\ntoken_secret=hf_secret.name for gated modelsAutomodelJobInput references use the workspace/name format: model=f"default/{MODEL_NAME}" and dataset={"training": f"default/{DATASET_NAME}"} (for example, default/llama-3-2-1b-base, default/sft-dataset)fileset=f"default/{MODEL_NAME}"client.jobs.get_status(name=job.job.name, workspace="default")Job fails with OOM (Out of Memory) error:
\nglobal_batch_size from 64 to 32 or 16 in batch={...}micro_batch_size at 1 (already the minimum in this tutorial)max_seq_length from 2048 to 1024 or 512 in training={...}num_gpus_per_node and tensor_parallel_size in parallelism={...}Loss curves not decreasing (underfitting):
\nepochs from 2 to 3-5 in schedule={...}1e-4 or 1e-5 instead of the default 5e-5 in optimizer={...}Training loss decreases but validation loss increases (overfitting):
\nepochs from 2 to 1 in schedule={...}learning_rate from 5e-5 to 2e-5 or 1e-5 in optimizer={...}Model output quality is poor despite good training metrics:
\nlearning_rate and global_batch_sizeDeployment fails:
\nclient.models.retrieve(name=OUTPUT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")executor_config={"gpu": 1, ...}engine="vllm" with vllm/vllm-openai:v0.22.1Manual Evaluation (Recommended)
\nWhat to look for:
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\nJob fails during model download:
\ntoken_secret=hf_secret.name for gated modelsAutomodelJobInput references use the workspace/name format: model=f"default/{MODEL_NAME}" and dataset={"training": f"default/{DATASET_NAME}"} (for example, default/llama-3-2-1b-base, default/sft-dataset)fileset=f"default/{MODEL_NAME}"client.jobs.get_status(name=job.job.name, workspace="default")Job fails with OOM (Out of Memory) error:
\nglobal_batch_size from 64 to 32 or 16 in batch={...}micro_batch_size at 1 (already the minimum in this tutorial)max_seq_length from 2048 to 1024 or 512 in training={...}num_gpus_per_node and tensor_parallel_size in parallelism={...}Loss curves not decreasing (underfitting):
\nepochs from 2 to 3-5 in schedule={...}1e-4 or 1e-5 instead of the default 5e-5 in optimizer={...}Training loss decreases but validation loss increases (overfitting):
\nepochs from 2 to 1 in schedule={...}learning_rate from 5e-5 to 2e-5 or 1e-5 in optimizer={...}Model output quality is poor despite good training metrics:
\nlearning_rate and global_batch_sizeDeployment fails:
\nclient.models.retrieve(name=OUTPUT_NAME, workspace="default")client.inference.deployments.get_logs(name=deployment.name, workspace="default")executor_config={"gpu": 1, ...}engine="vllm" with vllm/vllm-openai:v0.22.1