Create your data in JSONL format - one JSON object per line. The platform auto-detects your data format. Supported dataset formats are listed below.
\nThis format contains complete conversation histories for both the chosen (preferred) and rejected responses.
\nThis format uses numeric preference scores to indicate which response is better. The context can be either a simple string or an array of message objects.
\nIf you plan to use NGC or HuggingFace models, you'll need to configure authentication:
\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.
\n"
- },
- {
- "type": "markdown",
- "source": "**Target `model_uri` Format:**\n\nCurrently, `model_uri` must reference a FileSet:\n- **FileSet:** `fileset://{workspace}/{fileset-name}`\n\nSupport for direct HuggingFace (`hf://`) and NGC (`ngc://`) URIs is coming soon. For now, create a fileset as shown in the previous step, and the HuggingFace model will be downloaded at the beginning the finetuning job.\n\n**GPU Requirements:**\n- 1B models: 1 GPU (24GB+ VRAM)\n- 3B models: 1-2 GPUs \n- 8B models: 2-4 GPUs\n- 70B models: 8+ GPUs \n\nAdjust `num_gpus_per_node` and `tensor_parallel_size` based on your model size.\n\n**Important**\n\nWhen setting `val_check_interval` for DPO, use a fractional value (e.g., `0.5` for twice per epoch) or omit it entirely (validates once at end of epoch). Avoid integer step counts — they may not divide evenly into the total training steps, which can prevent validation from running on the final step.",
- "source_html": "For larger models requiring multiple GPUs, configure parallelism with environment variables:
\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the NIM_TENSOR_PARALLEL_SIZE and NIM_PIPELINE_PARALLEL_SIZE environment variables.
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, see the Hyperparameter Reference.
\nLearn how to use the NeMo Platform to create a DPO (Direct Preference Optimization) job using a custom dataset.
\nDPO is an advanced fine-tuning technique for preference-based alignment. If you're new to fine-tuning, consider starting with LoRA or Full SFT tutorials first.
\nDirect Preference Optimization (DPO) is an RL-free alignment algorithm that operates on preference data. Given a prompt and a pair of chosen and rejected responses, DPO aims to increase the probability of the chosen response and decrease the probability of the rejected response relative to a frozen reference model. The actor is initialized using the reference model. For more details, refer to the DPO paper.
\nDPO shares similarities with Full SFT training workflows but differs in a few key ways:
\n\n```",
- "source_html": "Quick Start
\n1. Initialize SDK
\nThe SDK needs to know your NeMo Platform server URL. By default, http://localhost:8080 is used in accordance with the Quickstart guide. If NeMo Platform is running at a custom location, you can override the URL by setting the NMP_BASE_URL environment variable:
\nexport NMP_BASE_URL=<YOUR_NMP_BASE_URL>\n
\n"
- },
- {
- "type": "code",
- "source": "import json\nimport os\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n DpoTrainingParam,\n ParallelismParamsParam,\n)\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace=\"default\"\n)",
- "language": "python",
- "source_html": "import json\nimport os\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n DpoTrainingParam,\n ParallelismParamsParam,\n)\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace="default"\n)\n"
- },
- {
- "type": "markdown",
- "source": "### 2. Prepare Dataset\n\nCreate your data in JSONL format - one JSON object per line. The platform auto-detects your data format. Supported dataset formats are listed below.\n\n**Flexible Data Setup:**\n- **No validation file?** The platform automatically creates a 10% validation split\n- **Multiple files?** Upload to `training/` or `validation/` subdirectories—they'll be automatically merged\n- **Format detection:** Your data format is auto-detected at training time\n\nIn this tutorial the following dataset directory structure will be used:\n```\nmy_dataset\n`-- training.jsonl\n`-- validation.jsonl\n```",
- "source_html": "2. Prepare Dataset
\nCreate your data in JSONL format - one JSON object per line. The platform auto-detects your data format. Supported dataset formats are listed below.
\nFlexible Data Setup:
\n\n- No validation file? The platform automatically creates a 10% validation split
\n- Multiple files? Upload to
training/ or validation/ subdirectories—they'll be automatically merged \n- Format detection: Your data format is auto-detected at training time
\n
\nIn this tutorial the following dataset directory structure will be used:
\nmy_dataset\n`-- training.jsonl\n`-- validation.jsonl\n
\n"
- },
- {
- "type": "markdown",
- "source": "#### Binary Preference Format\nDPO training requires preference pairs with three fields:\n- **`prompt`**: The input prompt (can be a string or array of message objects)\n- **`chosen`**: The preferred response\n- **`rejected`**: The less preferred response",
- "source_html": "Binary Preference Format
\nDPO training requires preference pairs with three fields:
\n\nprompt: The input prompt (can be a string or array of message objects) \nchosen: The preferred response \nrejected: The less preferred response \n
\n"
- },
- {
- "type": "code",
- "source": "{\"prompt\": [{\"role\": \"user\", \"content\": \"What is the capital of France?\"}], \"chosen\": \"The capital of France is Paris. It is the largest city in France and serves as the country's political, economic, and cultural center.\", \"rejected\": \"I think the capital of France might be London or Paris, I'm not entirely sure.\"}",
- "language": "python",
- "source_html": "{"prompt": [{"role": "user", "content": "What is the capital of France?"}], "chosen": "The capital of France is Paris. It is the largest city in France and serves as the country's political, economic, and cultural center.", "rejected": "I think the capital of France might be London or Paris, I'm not entirely sure."}\n"
- },
- {
- "type": "markdown",
- "source": "#### Tulu3 Preference Dataset Format\nThis format contains complete conversation histories for both the chosen (preferred) and rejected responses.\n\nRequired fields:\n- **`chosen`**: Full conversation with the preferred response (list of message objects, last must be assistant)\n- **`rejected`**: Full conversation with the rejected response (list of message objects, last must be assistant)",
- "source_html": "Tulu3 Preference Dataset Format
\nThis format contains complete conversation histories for both the chosen (preferred) and rejected responses.
\nRequired fields:
\n\nchosen: Full conversation with the preferred response (list of message objects, last must be assistant) \nrejected: Full conversation with the rejected response (list of message objects, last must be assistant) \n
\n"
- },
- {
- "type": "code",
- "source": "{\"chosen\": [{\"role\": \"user\", \"content\": \"What is the capital of France?\"}, {\"role\": \"assistant\", \"content\": \"The capital of France is Paris.\"}], \"rejected\": [{\"role\": \"user\", \"content\": \"What is the capital of France?\"}, {\"role\": \"assistant\", \"content\": \"I'm not sure, but I think it might be London or Paris.\"}]}",
- "language": "python",
- "source_html": "{"chosen": [{"role": "user", "content": "What is the capital of France?"}, {"role": "assistant", "content": "The capital of France is Paris."}], "rejected": [{"role": "user", "content": "What is the capital of France?"}, {"role": "assistant", "content": "I'm not sure, but I think it might be London or Paris."}]}\n"
- },
- {
- "type": "markdown",
- "source": "#### HelpSteer Dataset Format\nThis format uses numeric preference scores to indicate which response is better. The context can be either a simple string or an array of message objects.\n\nRequired fields:\n- **`context`**: The input context (can be a string or array of message objects)\n- **`response1`**: First response option\n- **`response2`**: Second response option\n- **`overall_preference`**: Preference score where negative values mean response1 is preferred, positive values mean response2 is preferred, and 0 indicates a tie",
- "source_html": "HelpSteer Dataset Format
\nThis format uses numeric preference scores to indicate which response is better. The context can be either a simple string or an array of message objects.
\nRequired fields:
\n\ncontext: The input context (can be a string or array of message objects) \nresponse1: First response option \nresponse2: Second response option \noverall_preference: Preference score where negative values mean response1 is preferred, positive values mean response2 is preferred, and 0 indicates a tie \n
\n"
- },
- {
- "type": "code",
- "source": "{\"context\": \"Explain how to use git rebase\", \"response1\": \"Git rebase is a command that rewrites commit history by moving or combining commits. Use 'git rebase main' to reapply your branch commits on top of main. This creates a linear history and avoids merge commits.\", \"response2\": \"Use git rebase to change commits. Just type git rebase and it will work.\", \"overall_preference\": -2}",
- "language": "python",
- "source_html": "{"context": "Explain how to use git rebase", "response1": "Git rebase is a command that rewrites commit history by moving or combining commits. Use 'git rebase main' to reapply your branch commits on top of main. This creates a linear history and avoids merge commits.", "response2": "Use git rebase to change commits. Just type git rebase and it will work.", "overall_preference": -2}\n"
- },
- {
- "type": "markdown",
- "source": "### 3. Create Dataset FileSet and Upload Training Data",
- "source_html": "3. Create Dataset FileSet and Upload Training Data
\n"
- },
- {
- "type": "markdown",
- "source": "Install huggingface datasets package to download public [nvidia/HelpSteer3](https://huggingface.co/datasets/nvidia/HelpSteer3) dataset if it's not installed in your Python environment:\n\n```sh\npip install datasets\n```",
- "source_html": "Install huggingface datasets package to download public nvidia/HelpSteer3 dataset if it's not installed in your Python environment:
\npip install datasets\n
\n"
- },
- {
- "type": "markdown",
- "source": "#### Download nvidia/HelpSteer3 Dataset",
- "source_html": "Download nvidia/HelpSteer3 Dataset
\n"
- },
- {
- "type": "code",
- "source": "from pathlib import Path\nfrom datasets import load_dataset, Dataset\nds = load_dataset(\"nvidia/HelpSteer3\", \"preference\")\n\n# Adjust these values to change the size of the training and validation sets\n# The larger the datasets, the better the model will perform but longer the training will take\n# For the purpose of this tutorial, we'll use a small subset of the dataset\ntraining_size = 3000\nvalidation_size = 300\nDATASET_PATH = Path(\"dpo-dataset\").absolute()\n\n# Get training split and verify it's a Dataset (not IterableDataset)\ntrain_dataset = ds[\"train\"]\nvalidation_dataset = ds[\"validation\"]\nassert isinstance(train_dataset, Dataset), \"Expected Dataset type\"\nassert isinstance(validation_dataset, Dataset), \"Expected Dataset type\"\n\n# Select subsets and save to JSONL files\ntesting_ds = train_dataset.select(range(training_size))\nvalidation_ds = validation_dataset.select(range(validation_size))\n\n# Create directory if it doesn't exist\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Save subsets to JSONL files\ntesting_ds.to_json(f\"{DATASET_PATH}/training.jsonl\")\nvalidation_ds.to_json(f\"{DATASET_PATH}/validation.jsonl\")\n\nprint(f\"Saved training.jsonl with {len(testing_ds)} rows\")\nprint(f\"Saved validation.jsonl with {len(validation_ds)} rows\")",
- "language": "python",
- "source_html": "from pathlib import Path\nfrom datasets import load_dataset, Dataset\nds = load_dataset("nvidia/HelpSteer3", "preference")\n\n# Adjust these values to change the size of the training and validation sets\n# The larger the datasets, the better the model will perform but longer the training will take\n# For the purpose of this tutorial, we'll use a small subset of the dataset\ntraining_size = 3000\nvalidation_size = 300\nDATASET_PATH = Path("dpo-dataset").absolute()\n\n# Get training split and verify it's a Dataset (not IterableDataset)\ntrain_dataset = ds["train"]\nvalidation_dataset = ds["validation"]\nassert isinstance(train_dataset, Dataset), "Expected Dataset type"\nassert isinstance(validation_dataset, Dataset), "Expected Dataset type"\n\n# Select subsets and save to JSONL files\ntesting_ds = train_dataset.select(range(training_size))\nvalidation_ds = validation_dataset.select(range(validation_size))\n\n# Create directory if it doesn't exist\nos.makedirs(DATASET_PATH, exist_ok=True)\n\n# Save subsets to JSONL files\ntesting_ds.to_json(f"{DATASET_PATH}/training.jsonl")\nvalidation_ds.to_json(f"{DATASET_PATH}/validation.jsonl")\n\nprint(f"Saved training.jsonl with {len(testing_ds)} rows")\nprint(f"Saved validation.jsonl with {len(validation_ds)} rows")\n"
- },
- {
- "type": "markdown",
- "source": "#### Upload Training Data",
- "source_html": "Upload Training Data
\n"
- },
- {
- "type": "code",
- "source": "# Create fileset to store DPO training data\nDATASET_NAME = \"dpo-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"dpo training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
- "language": "python",
- "source_html": "# Create fileset to store DPO training data\nDATASET_NAME = "dpo-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="dpo training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
- },
- {
- "type": "markdown",
- "source": "### 4. Secrets Setup\n\nIf you plan to use NGC or HuggingFace models, you'll need to configure authentication:\n\n- **NGC models** (`ngc://` URIs): Requires NGC API key\n- **HuggingFace models** (`hf://` URIs): Requires HF token for gated/private models\n\n\nConfigure these as secrets in your platform. See [Managing Secrets](../../get-started/concepts/manage-secrets.md) for detailed instructions.\n\nGet your credentials to access base models:\n- [NGC API Key](https://ngc.nvidia.com/) (Setup → Generate API Key)\n- [HuggingFace Token](https://huggingface.co/settings/tokens) (Create token with Read access)\n\n\n---\n\n#### Quick Setup Example\n\nIn this tutorial we are going to work with [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct/tree/main) model from HuggingFace. Ensure that you have sufficient permissions to download the model. If you cannot see the files in the [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct/tree/main) Hugging Face page, request access\n\n**HuggingFace Authentication:**\n- For gated models (Llama, Gemma), you must provide a HuggingFace token via the `token_secret` parameter\n- Get your token from [HuggingFace Settings](https://huggingface.co/settings/tokens) (requires Read access)\n- Accept the model's terms on the HuggingFace model page before using it. Example: [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct/tree/main)\n- For public models, you can omit the `token_secret` parameter when creating a fileset for model in the next step",
- "source_html": "4. Secrets Setup
\nIf you plan to use NGC or HuggingFace models, you'll need to configure authentication:
\n\n- NGC models (
ngc:// URIs): Requires NGC API key \n- HuggingFace models (
hf:// URIs): Requires HF token for gated/private models \n
\nConfigure these as secrets in your platform. See Managing Secrets for detailed instructions.
\nGet your credentials to access base models:
\n\n
\nQuick Setup Example
\nIn this tutorial we are going to work with meta-llama/Llama-3.2-1B-Instruct model from HuggingFace. Ensure that you have sufficient permissions to download the model. If you cannot see the files in the meta-llama/Llama-3.2-1B-Instruct Hugging Face page, request access
\nHuggingFace Authentication:
\n\n- For gated models (Llama, Gemma), you must provide a HuggingFace token via the
token_secret parameter \n- Get your token from HuggingFace Settings (requires Read access)
\n- Accept the model's terms on the HuggingFace model page before using it. Example: meta-llama/Llama-3.2-1B-Instruct
\n- For public models, you can omit the
token_secret parameter when creating a fileset for model in the next step \n
\n"
- },
- {
- "type": "code",
- "source": "# Export the HF_TOKEN and NGC_API_KEY environment variables if they are not already set\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\nNGC_API_KEY = os.getenv(\"NGC_API_KEY\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(\"HF_TOKEN secret:\")\nprint(hf_secret.model_dump_json(indent=2))\n\n# Create NGC API key secret\n# Uncomment the line below if you have NGC API Key and want to finetune NGC models\n# ngc_api_key = create_or_get_secret(\"ngc-api-key\", NGC_API_KEY, \"NGC_API_KEY\")",
- "language": "python",
- "source_html": "# Export the HF_TOKEN and NGC_API_KEY environment variables if they are not already set\nHF_TOKEN = os.getenv("HF_TOKEN")\nNGC_API_KEY = os.getenv("NGC_API_KEY")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint("HF_TOKEN secret:")\nprint(hf_secret.model_dump_json(indent=2))\n\n# Create NGC API key secret\n# Uncomment the line below if you have NGC API Key and want to finetune NGC models\n# ngc_api_key = create_or_get_secret("ngc-api-key", NGC_API_KEY, "NGC_API_KEY")\n"
- },
- {
- "type": "markdown",
- "source": "### 5. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct/tree/main) model in HuggingFace that we will train with DPO. Then create a Model Entity that references this fileset. Model downloading will take place at the DPO finetuning job creation time.\n\nNote: for public models, you can omit the `token_secret` parameter when creating a model fileset.",
- "source_html": "5. Create Base Model FileSet and Model Entity
\nCreate a fileset pointing to meta-llama/Llama-3.2-1B-Instruct model in HuggingFace that we will train with DPO. Then create a Model Entity that references this fileset. Model downloading will take place at the DPO finetuning job creation time.
\nNote: for public models, you can omit the token_secret parameter when creating a model fileset.
\n"
- },
- {
- "type": "code",
- "source": "import time\n\n# Create a fileset pointing to the desired HuggingFace model\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-base\"\n\n# Ensure you have a HuggingFace token secret created\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=\"default\",\n name=MODEL_NAME,\n description=\"Llama 3.2 1B base model from HuggingFace\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n # repo_id is the full model name from Hugging Face\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n # we use the secret created in the previous step\n token_secret=hf_secret.name\n )\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=\"default\",\n name=MODEL_NAME,\n fileset=f\"default/{MODEL_NAME}\",\n )\n\nprint(f\"\\nBase model fileset: fileset://default/{base_model.name}\")\nprint(\"Base model fileset files list:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=\"default\").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=\"default\",\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec}\")",
- "language": "python",
- "source_html": "import time\n\n# Create a fileset pointing to the desired HuggingFace model\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-base"\n\n# Ensure you have a HuggingFace token secret created\ntry:\n base_model_fs = client.files.filesets.create(\n workspace="default",\n name=MODEL_NAME,\n description="Llama 3.2 1B base model from HuggingFace",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n # repo_id is the full model name from Hugging Face\n repo_id=HF_REPO_ID,\n repo_type="model",\n # we use the secret created in the previous step\n token_secret=hf_secret.name\n )\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\n# Create Model Entity referencing the FileSet\ntry:\n base_model = client.models.create(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace="default",\n name=MODEL_NAME,\n fileset=f"default/{MODEL_NAME}",\n )\n\nprint(f"\\nBase model fileset: fileset://default/{base_model.name}")\nprint("Base model fileset files list:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace="default").data], indent=2))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace="default",\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec}")\n"
- },
- {
- "type": "markdown",
- "source": "### 6. Create DPO Finetuning Job\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.",
- "source_html": "6. Create DPO Finetuning Job
\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.
\n"
- },
- {
- "type": "markdown",
- "source": "**Target `model_uri` Format:**\n\nCurrently, `model_uri` must reference a FileSet:\n- **FileSet:** `fileset://{workspace}/{fileset-name}`\n\nSupport for direct HuggingFace (`hf://`) and NGC (`ngc://`) URIs is coming soon. For now, create a fileset as shown in the previous step, and the HuggingFace model will be downloaded at the beginning the finetuning job.\n\n**GPU Requirements:**\n- 1B models: 1 GPU (24GB+ VRAM)\n- 3B models: 1-2 GPUs \n- 8B models: 2-4 GPUs\n- 70B models: 8+ GPUs \n\nAdjust `num_gpus_per_node` and `tensor_parallel_size` based on your model size.\n\n**Important**\n\nWhen setting `val_check_interval` for DPO, use a fractional value (e.g., `0.5` for twice per epoch) or omit it entirely (validates once at end of epoch). Avoid integer step counts — they may not divide evenly into the total training steps, which can prevent validation from running on the final step.",
- "source_html": "Target model_uri Format:
\nCurrently, model_uri must reference a FileSet:
\n\n- FileSet:
fileset://{workspace}/{fileset-name} \n
\nSupport for direct HuggingFace (hf://) and NGC (ngc://) URIs is coming soon. For now, create a fileset as shown in the previous step, and the HuggingFace model will be downloaded at the beginning the finetuning job.
\nGPU Requirements:
\n\n- 1B models: 1 GPU (24GB+ VRAM)
\n- 3B models: 1-2 GPUs
\n- 8B models: 2-4 GPUs
\n- 70B models: 8+ GPUs
\n
\nAdjust num_gpus_per_node and tensor_parallel_size based on your model size.
\nImportant
\nWhen setting val_check_interval for DPO, use a fractional value (e.g., 0.5 for twice per epoch) or omit it entirely (validates once at end of epoch). Avoid integer step counts — they may not divide evenly into the total training steps, which can prevent validation from running on the final step.
\n"
- },
- {
- "type": "code",
- "source": "import uuid\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-dpo-job-{job_suffix}\"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=DpoTrainingParam(\n type=\"dpo\",\n epochs=1,\n batch_size=16,\n learning_rate=0.00005,\n max_seq_length=4096,\n ref_policy_kl_penalty=0.1,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n )\n)\n\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
- "language": "python",
- "source_html": "import uuid\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-dpo-job-{job_suffix}"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=DpoTrainingParam(\n type="dpo",\n epochs=1,\n batch_size=16,\n learning_rate=0.00005,\n max_seq_length=4096,\n ref_policy_kl_penalty=0.1,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n )\n)\n\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
- },
- {
- "type": "markdown",
- "source": "### 7. Track Training Progress",
- "source_html": "7. Track Training Progress
\n"
- },
- {
- "type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
- "language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
- },
- {
- "type": "markdown",
- "source": "**Interpreting DPO Training Metrics:**\n\nDPO training produces several key metrics:\n\n| Metric | Description | What to Look For |\n|--------|-------------|------------------|\n| **loss** | Total training loss (preference_loss + sft_loss) | Should decrease over training |\n| **preference_loss** | Core DPO loss measuring preference learning | Starts near ln(2) ≈ 0.693, should decrease |\n| **sft_loss** | SFT regularization term (often 0 for pure DPO) | Depends on configuration |\n| **accuracy** | Fraction of samples where chosen > rejected | Should increase toward 80-95%+ |\n| **rewards_chosen_mean** | Average implicit reward for chosen responses | Should be positive |\n| **rewards_rejected_mean** | Average implicit reward for rejected responses | Should be negative |\n\n**Key Indicators:**\n\n- **Reward Margin** = `rewards_chosen_mean - rewards_rejected_mean`\n - Should be positive and increasing\n - Indicates the model is learning to distinguish preferences\n\n- **Accuracy Interpretation:**\n - 50% = random chance (no learning)\n - 66-75% = early/moderate learning\n - 80%+ = good preference learning\n - 95%+ = strong preference alignment\n\n**Troubleshooting:**\n\n- **Loss near ln(2) ≈ 0.693**: Model is at random chance level, training just starting or not learning\n- **Accuracy stuck at ~50%**: Check data quality, increase learning rate, or verify preference labels\n- **Negative reward margin**: Model is learning the wrong direction—check chosen/rejected labels\n- **Loss increasing**: Learning rate too high or data quality issues\n\n**Note:** Training metrics measure optimization progress, not final model quality. Always evaluate the deployed model on your specific use case.",
- "source_html": "Interpreting DPO Training Metrics:
\nDPO training produces several key metrics:
\n\n\n\n| Metric | \nDescription | \nWhat to Look For | \n
\n\n\n\n| loss | \nTotal training loss (preference_loss + sft_loss) | \nShould decrease over training | \n
\n\n| preference_loss | \nCore DPO loss measuring preference learning | \nStarts near ln(2) ≈ 0.693, should decrease | \n
\n\n| sft_loss | \nSFT regularization term (often 0 for pure DPO) | \nDepends on configuration | \n
\n\n| accuracy | \nFraction of samples where chosen > rejected | \nShould increase toward 80-95%+ | \n
\n\n| rewards_chosen_mean | \nAverage implicit reward for chosen responses | \nShould be positive | \n
\n\n| rewards_rejected_mean | \nAverage implicit reward for rejected responses | \nShould be negative | \n
\n\n
\nKey Indicators:
\n\nTroubleshooting:
\n\n- Loss near ln(2) ≈ 0.693: Model is at random chance level, training just starting or not learning
\n- Accuracy stuck at ~50%: Check data quality, increase learning rate, or verify preference labels
\n- Negative reward margin: Model is learning the wrong direction—check chosen/rejected labels
\n- Loss increasing: Learning rate too high or data quality issues
\n
\nNote: Training metrics measure optimization progress, not final model quality. Always evaluate the deployed model on your specific use case.
\n"
- },
- {
- "type": "markdown",
- "source": "### 8. Deploy Fine-Tuned Model\n\nOnce training completes, deploy using the Deployment Management Service:",
- "source_html": "8. Deploy Fine-Tuned Model
\nOnce training completes, deploy using the Deployment Management Service:
\n"
- },
- {
- "type": "code",
- "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))",
- "language": "python",
- "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))\n"
- },
- {
- "type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"dpo-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"dpo-model-deployment-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=\"nvcr.io/nim/nvidia/llm-nim\",\n image_tag=\"1.13.1\",\n gpu=1,\n model_name=job.spec.output.name, # ModelEntity name from training,\n model_namespace=\"default\", # Workspace where ModelEntity lives\n )\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace=\"default\"\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {deployment_status.status}\")",
- "language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"dpo-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"dpo-model-deployment-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name="nvcr.io/nim/nvidia/llm-nim",\n image_tag="1.13.1",\n gpu=1,\n model_name=job.spec.output.name, # ModelEntity name from training,\n model_namespace="default", # Workspace where ModelEntity lives\n )\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace="default"\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {deployment_status.status}")\n"
- },
- {
- "type": "markdown",
- "source": "### Monitor status of deployment",
- "source_html": "Monitor status of deployment
\n"
- },
- {
- "type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll deployment status every 15 seconds until ready\nTIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f\"Monitoring deployment '{deployment.name}'...\")\nprint(f\"Timeout: {TIMEOUT_MINUTES} minutes\\n\")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace=\"default\"\n )\n \n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n \n clear_output(wait=True)\n print(f\"Deployment: {deployment.name}\")\n print(f\"Status: {deployment_status.status}\")\n print(f\"Elapsed time: {elapsed_min}m {elapsed_sec}s\")\n \n # Check if deployment is ready\n if deployment_status.status == \"READY\":\n print(\"\\nDeployment is ready!\")\n if not client.models.wait_for_gateway(deployment.name, workspace=\"default\", timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n \n # Check for failure states\n if deployment_status.status in (\"FAILED\", \"ERROR\", \"TERMINATED\", \"LOST\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n \n # Check timeout\n if elapsed > timeout_seconds:\n raise TimeoutError(f\"Deployment timeout after {TIMEOUT_MINUTES} minutes\")\n \n time.sleep(15)",
- "language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll deployment status every 15 seconds until ready\nTIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f"Monitoring deployment '{deployment.name}'...")\nprint(f"Timeout: {TIMEOUT_MINUTES} minutes\\n")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace="default"\n )\n \n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n \n clear_output(wait=True)\n print(f"Deployment: {deployment.name}")\n print(f"Status: {deployment_status.status}")\n print(f"Elapsed time: {elapsed_min}m {elapsed_sec}s")\n \n # Check if deployment is ready\n if deployment_status.status == "READY":\n print("\\nDeployment is ready!")\n if not client.models.wait_for_gateway(deployment.name, workspace="default", timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n \n # Check for failure states\n if deployment_status.status in ("FAILED", "ERROR", "TERMINATED", "LOST"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n \n # Check timeout\n if elapsed > timeout_seconds:\n raise TimeoutError(f"Deployment timeout after {TIMEOUT_MINUTES} minutes")\n \n time.sleep(15)\n"
- },
- {
- "type": "markdown",
- "source": "The deployment service automatically:\n- Downloads model weights from the Files service\n- Provisions storage (PVC) for the weights\n- Configures and starts the NIM container\n\n**Multi-GPU Deployment:**\n\nFor larger models requiring multiple GPUs, configure parallelism with environment variables:\n\n```python\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=\"sft-model-config-multigpu\",\n \n nim_deployment={\n \"image_name\": \"nvcr.io/nim/nvidia/llm-nim\",\n \"image_tag\": \"1.13.1\",\n \"gpu\": 2, # Total GPUs\n \"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"2\", # Tensor parallelism\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"1\" # Pipeline parallelism\n }\n }\n)\n```",
- "source_html": "The deployment service automatically:
\n\n- Downloads model weights from the Files service
\n- Provisions storage (PVC) for the weights
\n- Configures and starts the NIM container
\n
\nMulti-GPU Deployment:
\nFor larger models requiring multiple GPUs, configure parallelism with environment variables:
\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name="sft-model-config-multigpu",\n \n nim_deployment={\n "image_name": "nvcr.io/nim/nvidia/llm-nim",\n "image_tag": "1.13.1",\n "gpu": 2, # Total GPUs\n "additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "2", # Tensor parallelism\n "NIM_PIPELINE_PARALLEL_SIZE": "1" # Pipeline parallelism\n }\n }\n)\n
\n"
- },
- {
- "type": "markdown",
- "source": "**Single-Node Constraint:** Model deployments are limited to a single node. The maximum `gpu` value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.\n\n---\n\n#### GPU Parallelism\n\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the `NIM_TENSOR_PARALLEL_SIZE` and `NIM_PIPELINE_PARALLEL_SIZE` environment variables.\n\n| Strategy | Description | Best For |\n|----------|-------------|----------|\n| **Tensor Parallel (TP)** | Splits model layers across GPUs | Lowest latency |\n| **Pipeline Parallel (PP)** | Splits model depth across GPUs | Highest throughput |\n\n**Formula:** `gpu` = `NIM_TENSOR_PARALLEL_SIZE` × `NIM_PIPELINE_PARALLEL_SIZE`\n\n---\n\n#### Example Configurations\n\n**Default (TP=8, PP=1) — Lowest Latency**\n```\n\"gpu\": 8\n# NIM automatically sets NIM_TENSOR_PARALLEL_SIZE=8\n```\n\n**Balanced (TP=4, PP=2)**\n```\n\"gpu\": 8,\n\"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"4\",\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"2\"\n}\n```\n\n**Throughput Optimized (TP=2, PP=4)**\n```\n\"gpu\": 8,\n\"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"2\",\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"4\"\n}\n```",
- "source_html": "Single-Node Constraint: Model deployments are limited to a single node. The maximum gpu value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.
\n
\nGPU Parallelism
\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the NIM_TENSOR_PARALLEL_SIZE and NIM_PIPELINE_PARALLEL_SIZE environment variables.
\n\n\n\n| Strategy | \nDescription | \nBest For | \n
\n\n\n\n| Tensor Parallel (TP) | \nSplits model layers across GPUs | \nLowest latency | \n
\n\n| Pipeline Parallel (PP) | \nSplits model depth across GPUs | \nHighest throughput | \n
\n\n
\nFormula: gpu = NIM_TENSOR_PARALLEL_SIZE × NIM_PIPELINE_PARALLEL_SIZE
\n
\nExample Configurations
\nDefault (TP=8, PP=1) — Lowest Latency
\n"gpu": 8\n# NIM automatically sets NIM_TENSOR_PARALLEL_SIZE=8\n
\nBalanced (TP=4, PP=2)
\n"gpu": 8,\n"additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "4",\n "NIM_PIPELINE_PARALLEL_SIZE": "2"\n}\n
\nThroughput Optimized (TP=2, PP=4)
\n"gpu": 8,\n"additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "2",\n "NIM_PIPELINE_PARALLEL_SIZE": "4"\n}\n
\n"
- },
- {
- "type": "markdown",
- "source": "### 9. Evaluate Your Model\n\nAfter training, evaluate whether your model meets your requirements:\n\n#### Quick Manual Evaluation",
- "source_html": "9. Evaluate Your Model
\nAfter training, evaluate whether your model meets your requirements:
\nQuick Manual Evaluation
\n"
- },
- {
- "type": "code",
- "source": "# Wait for deployment to be ready, then test\nmessages = [\n {\"role\": \"system\", \"content\": \"You are a helpful assistant.\"},\n {\"role\": \"user\", \"content\": \"Write a short email to my colleague.\"}\n]\n\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": f\"default/{job.spec.output.name}\", # Match the model_name from deployment config\n \"messages\": messages,\n \"temperature\": 0.7,\n \"max_tokens\": 256\n }\n)\n\n# Display prompt and completion\nprint(\"=\" * 60)\nprint(\"PROMPT\")\nprint(\"=\" * 60)\nfor msg in messages:\n print(f\"[{msg['role'].upper()}]\")\n print(msg[\"content\"])\n print()\n\nprint(\"=\" * 60)\nprint(\"COMPLETION\")\nprint(\"=\" * 60)\nprint(\"[ASSISTANT]\")\ncompletion = response[\"choices\"][0][\"message\"][\"content\"]\nprint(completion)",
- "language": "python",
- "source_html": "# Wait for deployment to be ready, then test\nmessages = [\n {"role": "system", "content": "You are a helpful assistant."},\n {"role": "user", "content": "Write a short email to my colleague."}\n]\n\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={\n "model": f"default/{job.spec.output.name}", # Match the model_name from deployment config\n "messages": messages,\n "temperature": 0.7,\n "max_tokens": 256\n }\n)\n\n# Display prompt and completion\nprint("=" * 60)\nprint("PROMPT")\nprint("=" * 60)\nfor msg in messages:\n print(f"[{msg['role'].upper()}]")\n print(msg["content"])\n print()\n\nprint("=" * 60)\nprint("COMPLETION")\nprint("=" * 60)\nprint("[ASSISTANT]")\ncompletion = response["choices"][0]["message"]["content"]\nprint(completion)\n"
- },
- {
- "type": "markdown",
- "source": "#### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world examples from your use case\n- Compare responses to base model and expected outputs\n- Verify the model exhibits desired behavior changes\n- Check edge cases and error handling\n\n**What to look for:**\n- ✅ Model follows your desired output format\n- ✅ Applies domain knowledge correctly\n- ✅ Maintains general language capabilities\n- ✅ Avoids unwanted behaviors or biases\n- ❌ Doesn't hallucinate facts not in training data\n- ❌ Doesn't produce repetitive or nonsensical outputs\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, see the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n---\n\n\n## Troubleshooting\n\n**Job fails during model download:**\n- Verify authentication secrets are configured (see [Managing Secrets](../../get-started/concepts/manage-secrets.md))\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page\n- Check the `model_uri` format is correct (`fileset://`)\n- Ensure you have accepted the model's terms of service on HuggingFace\n- Check job status and logs: `client.customization.jobs.retrieve(name=job.name, workspace=\"default\")`\n\n**Job fails with OOM (Out of Memory) error:**\n1. **First try:** Reduce `micro_batch_size` from 2 to 1\n2. **Still OOM:** Reduce `batch_size` from 16 to 8\n3. **Still OOM:** Reduce `max_seq_length` from 2048 to 1024 or 512\n4. **Last resort:** Increase GPU count and use `tensor_parallel_size` for model sharding\n\n**Loss curves not decreasing (underfitting):**\n- Increase training duration: `epochs: 5-10` instead of 3\n- Adjust learning rate: Try `1e-5` to `1e-4`\n- Add warmup: Set `warmup_steps` to ~10% of total training steps\n- Check data quality: Verify formatting, remove duplicates, ensure diversity\n\n**Training loss decreases but validation loss increases (overfitting):**\n- Reduce epochs: Try `epochs: 1-2` instead of 5+\n- Lower learning rate: Use `2e-5` or `1e-5`\n- Increase dataset size and diversity\n- Verify train/validation split has no data leakage\n\n**Model output quality is poor despite good training metrics:**\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases\n- Review data quality, format, and diversity—metrics can be misleading with poor data\n- Try a different base model size or architecture\n- Adjust learning rate and batch size\n- Compare to baseline: Test base model to ensure fine-tuning improved performance\n\n**Deployment fails:**\n- Verify output model exists: `client.models.retrieve(name=job.spec.output.name, workspace=\"default\")`\n- Check deployment logs: `client.inference.deployments.get_logs(name=deployment.name, workspace=\"default\")`\n- Ensure sufficient GPU resources available for model size\n- Verify NIM image tag `1.13.1` is compatible with your model\n\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Learn about [LoRA customization](./lora-customization-job) for resource-efficient fine-tuning",
- "source_html": "Evaluation Best Practices
\nManual Evaluation (Recommended)
\n\n- Test with real-world examples from your use case
\n- Compare responses to base model and expected outputs
\n- Verify the model exhibits desired behavior changes
\n- Check edge cases and error handling
\n
\nWhat to look for:
\n\n- ✅ Model follows your desired output format
\n- ✅ Applies domain knowledge correctly
\n- ✅ Maintains general language capabilities
\n- ✅ Avoids unwanted behaviors or biases
\n- ❌ Doesn't hallucinate facts not in training data
\n- ❌ Doesn't produce repetitive or nonsensical outputs
\n
\n
\nHyperparameters
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, see the Hyperparameter Reference.
\n
\nTroubleshooting
\nJob fails during model download:
\n\n- Verify authentication secrets are configured (see Managing Secrets)
\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page
\n- Check the
model_uri format is correct (fileset://) \n- Ensure you have accepted the model's terms of service on HuggingFace
\n- Check job status and logs:
client.customization.jobs.retrieve(name=job.name, workspace="default") \n
\nJob fails with OOM (Out of Memory) error:
\n\n- First try: Reduce
micro_batch_size from 2 to 1 \n- Still OOM: Reduce
batch_size from 16 to 8 \n- Still OOM: Reduce
max_seq_length from 2048 to 1024 or 512 \n- Last resort: Increase GPU count and use
tensor_parallel_size for model sharding \n
\nLoss curves not decreasing (underfitting):
\n\n- Increase training duration:
epochs: 5-10 instead of 3 \n- Adjust learning rate: Try
1e-5 to 1e-4 \n- Add warmup: Set
warmup_steps to ~10% of total training steps \n- Check data quality: Verify formatting, remove duplicates, ensure diversity
\n
\nTraining loss decreases but validation loss increases (overfitting):
\n\n- Reduce epochs: Try
epochs: 1-2 instead of 5+ \n- Lower learning rate: Use
2e-5 or 1e-5 \n- Increase dataset size and diversity
\n- Verify train/validation split has no data leakage
\n
\nModel output quality is poor despite good training metrics:
\n\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases
\n- Review data quality, format, and diversity—metrics can be misleading with poor data
\n- Try a different base model size or architecture
\n- Adjust learning rate and batch size
\n- Compare to baseline: Test base model to ensure fine-tuning improved performance
\n
\nDeployment fails:
\n\n- Verify output model exists:
client.models.retrieve(name=job.spec.output.name, workspace="default") \n- Check deployment logs:
client.inference.deployments.get_logs(name=deployment.name, workspace="default") \n- Ensure sufficient GPU resources available for model size
\n- Verify NIM image tag
1.13.1 is compatible with your model \n
\nNext Steps
\n\n"
- }
-] };
diff --git a/docs/fern/components/notebooks/embedding-customization-job.json b/docs/fern/components/notebooks/embedding-customization-job.json
index 21b42099e2..0f4e98cd95 100644
--- a/docs/fern/components/notebooks/embedding-customization-job.json
+++ b/docs/fern/components/notebooks/embedding-customization-job.json
@@ -7,8 +7,8 @@
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)\n3. **HuggingFace token** with read access to download the SPECTER dataset (get one at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens))\n4. **NGC API key** to pull NIM container images from nvcr.io (get one at [ngc.nvidia.com](https://ngc.nvidia.com/) → Setup → Generate API Key)",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n- HuggingFace token with read access to download the SPECTER dataset (get one at huggingface.co/settings/tokens)
\n- NGC API key to pull NIM container images from nvcr.io (get one at ngc.nvidia.com → Setup → Generate API Key)
\n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **HuggingFace token** with read access to download the SPECTER dataset (get one at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens))\n4. **NGC API key** to pull NIM container images from nvcr.io (get one at [ngc.nvidia.com](https://ngc.nvidia.com/) → Setup → Generate API Key)",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n- HuggingFace token with read access to download the SPECTER dataset (get one at huggingface.co/settings/tokens)
\n- NGC API key to pull NIM container images from nvcr.io (get one at ngc.nvidia.com → Setup → Generate API Key)
\n
\n"
},
{
"type": "markdown",
@@ -40,9 +40,9 @@
},
{
"type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get(\"NGC_API_KEY\")\nif not NGC_API_KEY:\n raise ValueError(\"NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key\")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = \"ngc-api-key\"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace=\"default\", value=NGC_API_KEY)\n print(f\"Created secret: {NGC_SECRET_NAME}\")\nexcept ConflictError:\n print(f\"Secret '{NGC_SECRET_NAME}' already exists, continuing...\")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = \"nvidia/llama-nemotron-embed-1b-v2\"\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f\"baseline-embedding-cfg-{baseline_suffix}\"\nBASELINE_DEPLOYMENT_NAME = f\"baseline-embedding-{baseline_suffix}\"\n\nprint(\"Creating baseline deployment config...\")\nbaseline_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_CONFIG,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n image_pull_secret=NGC_SECRET_NAME,\n )\n)\n\nprint(\"Deploying base model...\")\nbaseline_deployment = client.inference.deployments.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f\"Baseline deployment: {baseline_deployment.name}\")",
+ "source": "# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get(\"NGC_API_KEY\")\nif not NGC_API_KEY:\n raise ValueError(\"NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key\")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = \"ngc-api-key\"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace=\"default\", value=NGC_API_KEY)\n print(f\"Created secret: {NGC_SECRET_NAME}\")\nexcept ConflictError:\n print(f\"Secret '{NGC_SECRET_NAME}' already exists, continuing...\")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = \"nvidia/llama-nemotron-embed-1b-v2\"\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f\"baseline-embedding-cfg-{baseline_suffix}\"\nBASELINE_DEPLOYMENT_NAME = f\"baseline-embedding-{baseline_suffix}\"\n\nprint(\"Creating baseline deployment config...\")\nbaseline_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_CONFIG,\n engine=\"nim\",\n model_spec={},\n executor_config={\n \"gpu\": 1,\n \"image_name\": NIM_IMAGE,\n \"image_tag\": NIM_TAG,\n },\n)\n\nprint(\"Deploying base model...\")\nbaseline_deployment = client.inference.deployments.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f\"Baseline deployment: {baseline_deployment.name}\")",
"language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get("NGC_API_KEY")\nif not NGC_API_KEY:\n raise ValueError("NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = "ngc-api-key"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace="default", value=NGC_API_KEY)\n print(f"Created secret: {NGC_SECRET_NAME}")\nexcept ConflictError:\n print(f"Secret '{NGC_SECRET_NAME}' already exists, continuing...")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = "nvidia/llama-nemotron-embed-1b-v2"\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f"baseline-embedding-cfg-{baseline_suffix}"\nBASELINE_DEPLOYMENT_NAME = f"baseline-embedding-{baseline_suffix}"\n\nprint("Creating baseline deployment config...")\nbaseline_config = client.inference.deployment_configs.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_CONFIG,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n image_pull_secret=NGC_SECRET_NAME,\n )\n)\n\nprint("Deploying base model...")\nbaseline_deployment = client.inference.deployments.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f"Baseline deployment: {baseline_deployment.name}")\n"
+ "source_html": "# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get("NGC_API_KEY")\nif not NGC_API_KEY:\n raise ValueError("NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = "ngc-api-key"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace="default", value=NGC_API_KEY)\n print(f"Created secret: {NGC_SECRET_NAME}")\nexcept ConflictError:\n print(f"Secret '{NGC_SECRET_NAME}' already exists, continuing...")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = "nvidia/llama-nemotron-embed-1b-v2"\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f"baseline-embedding-cfg-{baseline_suffix}"\nBASELINE_DEPLOYMENT_NAME = f"baseline-embedding-{baseline_suffix}"\n\nprint("Creating baseline deployment config...")\nbaseline_config = client.inference.deployment_configs.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_CONFIG,\n engine="nim",\n model_spec={},\n executor_config={\n "gpu": 1,\n "image_name": NIM_IMAGE,\n "image_tag": NIM_TAG,\n },\n)\n\nprint("Deploying base model...")\nbaseline_deployment = client.inference.deployments.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f"Baseline deployment: {baseline_deployment.name}")\n"
},
{
"type": "code",
@@ -85,9 +85,9 @@
},
{
"type": "code",
- "source": "# Create fileset to store embedding training data\nDATASET_NAME = \"embedding-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SPECTER embedding training data (scientific paper triplets)\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate upload\nprint(\"\\nUploaded files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
+ "source": "# Create fileset to store embedding training data\nDATASET_NAME = \"embedding-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SPECTER embedding training data (scientific paper triplets)\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files\nclient.files.upload(\n local_path=f\"{DATASET_PATH}/\",\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate upload\nprint(\"\\nUploaded files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
"language": "python",
- "source_html": "# Create fileset to store embedding training data\nDATASET_NAME = "embedding-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SPECTER embedding training data (scientific paper triplets)"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate upload\nprint("\\nUploaded files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
+ "source_html": "# Create fileset to store embedding training data\nDATASET_NAME = "embedding-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SPECTER embedding training data (scientific paper triplets)"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files\nclient.files.upload(\n local_path=f"{DATASET_PATH}/",\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate upload\nprint("\\nUploaded files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
},
{
"type": "markdown",
@@ -96,14 +96,14 @@
},
{
"type": "code",
- "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME} (created in Step 2)\")",
+ "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif \"NGC_SECRET_NAME\" not in globals():\n NGC_SECRET_NAME = \"ngc-api-key\"\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME}\")",
"language": "python",
- "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME} (created in Step 2)")\n"
+ "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif "NGC_SECRET_NAME" not in globals():\n NGC_SECRET_NAME = "ngc-api-key"\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME}")\n"
},
{
"type": "markdown",
- "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-3.2-nv-embedqa-1b-v2](https://huggingface.co/nvidia/llama-3.2-nv-embedqa-1b-v2) embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.\n\n---\n*Note*: Either `MODEL_NAME` or `HF_REPO_ID` below must contain the substring embed to indicate that this is an embedding model.",
- "source_html": "7. Create Base Model FileSet and Model Entity
\nCreate a fileset pointing to the nvidia/llama-3.2-nv-embedqa-1b-v2 embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n
\nNote: Either MODEL_NAME or HF_REPO_ID below must contain the substring embed to indicate that this is an embedding model.
\n"
+ "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-3.2-nv-embedqa-1b-v2](https://huggingface.co/nvidia/llama-3.2-nv-embedqa-1b-v2) embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.",
+ "source_html": "7. Create Base Model FileSet and Model Entity
\nCreate a fileset pointing to the nvidia/llama-3.2-nv-embedqa-1b-v2 embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n"
},
{
"type": "code",
@@ -113,14 +113,14 @@
},
{
"type": "markdown",
- "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training_type`**: `sft` (supervised fine-tuning)\n- **Full fine-tuning**: No `peft` config needed (omit for all-weights training)\n- **`learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook creates a job with all-weights finetuning but you can also run LoRA with `merge=True`, which trains a LoRA adapter and then merges it back into the base model after training. The final output is a standard full-weight checkpoint, identical in format to an all-weights fine-tuned model, but LoRA training is faster, uses less memory, and is more lenient in hyperparameter tuning.\n\nTo do that update the job request like so\n\n```\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n peft=LoRaParamsParam(\n type= \"lora\",\n merge=True\n )\n )\n )\n)\n```",
- "source_html": "8. Create Embedding Fine-tuning Job
\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nKey hyperparameters for embedding fine-tuning:
\n\ntraining_type: sft (supervised fine-tuning) \n- Full fine-tuning: No
peft config needed (omit for all-weights training) \nlearning_rate: Lower values (1e-6 to 5e-6) work well for embedding models \nbatch_size: Larger batches improve contrastive learning (128-256 recommended) \n
\nNOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook creates a job with all-weights finetuning but you can also run LoRA with merge=True, which trains a LoRA adapter and then merges it back into the base model after training. The final output is a standard full-weight checkpoint, identical in format to an all-weights fine-tuned model, but LoRA training is faster, uses less memory, and is more lenient in hyperparameter tuning.
\nTo do that update the job request like so
\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n peft=LoRaParamsParam(\n type= "lora",\n merge=True\n )\n )\n )\n)\n
\n"
+ "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with split `schedule`, `batch`, `optimizer`, and `parallelism` sections. Reference the model entity and dataset fileset by workspace/name (not `fileset://` URIs).\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training.training_type`**: `sft`\n- **`training.finetuning_type`**: `all_weights` for full fine-tuning, or `lora_merged` for merged LoRA\n- **`optimizer.learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch.global_batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set `finetuning_type` to `lora_merged`:\n\n```python\ntraining={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora_merged\",\n \"lora\": {\"rank\": 16, \"alpha\": 32},\n \"max_seq_length\": MAX_SEQ_LENGTH,\n}\n```",
+ "source_html": "8. Create Embedding Fine-tuning Job
\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nSubmit to the Automodel backend using AutomodelJobInput with split schedule, batch, optimizer, and parallelism sections. Reference the model entity and dataset fileset by workspace/name (not fileset:// URIs).
\nKey hyperparameters for embedding fine-tuning:
\n\ntraining.training_type: sft \ntraining.finetuning_type: all_weights for full fine-tuning, or lora_merged for merged LoRA \noptimizer.learning_rate: Lower values (1e-6 to 5e-6) work well for embedding models \nbatch.global_batch_size: Larger batches improve contrastive learning (128-256 recommended) \n
\nNOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set finetuning_type to lora_merged:
\ntraining={\n "training_type": "sft",\n "finetuning_type": "lora_merged",\n "lora": {"rank": 16, "alpha": 32},\n "max_seq_length": MAX_SEQ_LENGTH,\n}\n
\n"
},
{
"type": "code",
- "source": "from nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"embedding-finetune-job-{job_suffix}\"\n\n# Hyperparameters optimized for embedding fine-tuning\nEPOCHS = 1\nBATCH_SIZE = 128 # Larger batches help contrastive learning\nLEARNING_RATE = 5e-6 # Lower LR for embedding models\nMAX_SEQ_LENGTH = 512 # Typical for embedding models\n\n# Note: The 'name' field must contain 'embed' for the customizer to detect this as an embedding model\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n )\n)\n\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"embedding-finetune-job-{job_suffix}\"\nOUTPUT_NAME = f\"nv-embed-finetuned-{job_suffix}\"\n\nEPOCHS = 1\nBATCH_SIZE = 128\nLEARNING_RATE = 5e-6\nMAX_SEQ_LENGTH = 512\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"all_weights\",\n \"max_seq_length\": MAX_SEQ_LENGTH,\n },\n schedule={\"epochs\": EPOCHS},\n batch={\"global_batch_size\": BATCH_SIZE, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": LEARNING_RATE},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output model: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "from nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"embedding-finetune-job-{job_suffix}"\n\n# Hyperparameters optimized for embedding fine-tuning\nEPOCHS = 1\nBATCH_SIZE = 128 # Larger batches help contrastive learning\nLEARNING_RATE = 5e-6 # Lower LR for embedding models\nMAX_SEQ_LENGTH = 512 # Typical for embedding models\n\n# Note: The 'name' field must contain 'embed' for the customizer to detect this as an embedding model\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n )\n)\n\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"embedding-finetune-job-{job_suffix}"\nOUTPUT_NAME = f"nv-embed-finetuned-{job_suffix}"\n\nEPOCHS = 1\nBATCH_SIZE = 128\nLEARNING_RATE = 5e-6\nMAX_SEQ_LENGTH = 512\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "all_weights",\n "max_seq_length": MAX_SEQ_LENGTH,\n },\n schedule={"epochs": EPOCHS},\n batch={"global_batch_size": BATCH_SIZE, "micro_batch_size": 1},\n optimizer={"learning_rate": LEARNING_RATE},\n parallelism={"num_gpus_per_node": 1},\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output model: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -129,9 +129,9 @@
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
+ "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
},
{
"type": "markdown",
@@ -145,15 +145,15 @@
},
{
"type": "code",
- "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace=\"default\", name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))",
+ "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace=\"default\", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))",
"language": "python",
- "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace="default", name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))\n"
+ "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace="default", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))\n"
},
{
"type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"embedding-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"embedding-model-deployment-{deploy_suffix}\"\n\n# Embedding NIM image\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n model_name=job.spec.output.name,\n model_namespace=\"default\",\n )\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}\")",
+ "source": "# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"embedding-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"embedding-model-deployment-{deploy_suffix}\"\n\n# Embedding NIM image\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"nim\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": OUTPUT_NAME,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": NIM_IMAGE,\n \"image_tag\": NIM_TAG,\n },\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}\")",
"language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"embedding-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"embedding-model-deployment-{deploy_suffix}"\n\n# Embedding NIM image\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n model_name=job.spec.output.name,\n model_namespace="default",\n )\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}")\n"
+ "source_html": "# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"embedding-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"embedding-model-deployment-{deploy_suffix}"\n\n# Embedding NIM image\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="nim",\n model_spec={\n "model_namespace": "default",\n "model_name": OUTPUT_NAME,\n },\n executor_config={\n "gpu": 1,\n "image_name": NIM_IMAGE,\n "image_tag": NIM_TAG,\n },\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}")\n"
},
{
"type": "markdown",
@@ -173,9 +173,9 @@
},
{
"type": "code",
- "source": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f\"default/{job.spec.output.name}\"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": [DEMO_QUERY],\n \"input_type\": \"query\"\n }\n)\nquery_embedding = query_response[\"data\"][0][\"embedding\"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": DEMO_DOCS,\n \"input_type\": \"passage\"\n }\n)\ndoc_embeddings = [d[\"embedding\"] for d in doc_response[\"data\"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f\"Query: \\\"{DEMO_QUERY}\\\"\\n\")\nprint(f\"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}\")\nprint(\"-\" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f\"{DEMO_LABELS[b_idx]} [{b_score:.3f}]\" + (\" *\" if b_idx in DEMO_RELEVANT else \"\")\n f_label = f\"{DEMO_LABELS[f_idx]} [{f_score:.3f}]\" + (\" *\" if f_idx in DEMO_RELEVANT else \"\")\n \n print(f\"#{rank+1:<5} {b_label:<30} {f_label:<30}\")\n\nprint(\"\\n* = relevant paper\")\nprint(\"\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.\")",
+ "source": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f\"default/{OUTPUT_NAME}\"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": [DEMO_QUERY],\n \"input_type\": \"query\"\n }\n)\nquery_embedding = query_response[\"data\"][0][\"embedding\"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": DEMO_DOCS,\n \"input_type\": \"passage\"\n }\n)\ndoc_embeddings = [d[\"embedding\"] for d in doc_response[\"data\"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f\"Query: \\\"{DEMO_QUERY}\\\"\\n\")\nprint(f\"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}\")\nprint(\"-\" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f\"{DEMO_LABELS[b_idx]} [{b_score:.3f}]\" + (\" *\" if b_idx in DEMO_RELEVANT else \"\")\n f_label = f\"{DEMO_LABELS[f_idx]} [{f_score:.3f}]\" + (\" *\" if f_idx in DEMO_RELEVANT else \"\")\n \n print(f\"#{rank+1:<5} {b_label:<30} {f_label:<30}\")\n\nprint(\"\\n* = relevant paper\")\nprint(\"\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.\")",
"language": "python",
- "source_html": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f"default/{job.spec.output.name}"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": [DEMO_QUERY],\n "input_type": "query"\n }\n)\nquery_embedding = query_response["data"][0]["embedding"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": DEMO_DOCS,\n "input_type": "passage"\n }\n)\ndoc_embeddings = [d["embedding"] for d in doc_response["data"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f"Query: \\"{DEMO_QUERY}\\"\\n")\nprint(f"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}")\nprint("-" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f"{DEMO_LABELS[b_idx]} [{b_score:.3f}]" + (" *" if b_idx in DEMO_RELEVANT else "")\n f_label = f"{DEMO_LABELS[f_idx]} [{f_score:.3f}]" + (" *" if f_idx in DEMO_RELEVANT else "")\n \n print(f"#{rank+1:<5} {b_label:<30} {f_label:<30}")\n\nprint("\\n* = relevant paper")\nprint("\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.")\n"
+ "source_html": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f"default/{OUTPUT_NAME}"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": [DEMO_QUERY],\n "input_type": "query"\n }\n)\nquery_embedding = query_response["data"][0]["embedding"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": DEMO_DOCS,\n "input_type": "passage"\n }\n)\ndoc_embeddings = [d["embedding"] for d in doc_response["data"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f"Query: \\"{DEMO_QUERY}\\"\\n")\nprint(f"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}")\nprint("-" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f"{DEMO_LABELS[b_idx]} [{b_score:.3f}]" + (" *" if b_idx in DEMO_RELEVANT else "")\n f_label = f"{DEMO_LABELS[f_idx]} [{f_score:.3f}]" + (" *" if f_idx in DEMO_RELEVANT else "")\n \n print(f"#{rank+1:<5} {b_label:<30} {f_label:<30}")\n\nprint("\\n* = relevant paper")\nprint("\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.")\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/embedding-customization-job.ts b/docs/fern/components/notebooks/embedding-customization-job.ts
index 36cf919e9f..5b5a3aab41 100644
--- a/docs/fern/components/notebooks/embedding-customization-job.ts
+++ b/docs/fern/components/notebooks/embedding-customization-job.ts
@@ -1,7 +1,9 @@
-// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
-// SPDX-License-Identifier: Apache-2.0
-
-/** Auto-generated by ipynb-to-fern-json.py - do not edit */
+/**
+ * SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+ * SPDX-License-Identifier: Apache-2.0
+ *
+ * Auto-generated by ipynb-to-fern-json.py - do not edit manually.
+ */
export default { cells: [
{
"type": "markdown",
@@ -10,8 +12,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)\n3. **HuggingFace token** with read access to download the SPECTER dataset (get one at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens))\n4. **NGC API key** to pull NIM container images from nvcr.io (get one at [ngc.nvidia.com](https://ngc.nvidia.com/) → Setup → Generate API Key)",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n- HuggingFace token with read access to download the SPECTER dataset (get one at huggingface.co/settings/tokens)
\n- NGC API key to pull NIM container images from nvcr.io (get one at ngc.nvidia.com → Setup → Generate API Key)
\n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **HuggingFace token** with read access to download the SPECTER dataset (get one at [huggingface.co/settings/tokens](https://huggingface.co/settings/tokens))\n4. **NGC API key** to pull NIM container images from nvcr.io (get one at [ngc.nvidia.com](https://ngc.nvidia.com/) → Setup → Generate API Key)",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n- HuggingFace token with read access to download the SPECTER dataset (get one at huggingface.co/settings/tokens)
\n- NGC API key to pull NIM container images from nvcr.io (get one at ngc.nvidia.com → Setup → Generate API Key)
\n
\n"
},
{
"type": "markdown",
@@ -43,9 +45,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get(\"NGC_API_KEY\")\nif not NGC_API_KEY:\n raise ValueError(\"NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key\")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = \"ngc-api-key\"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace=\"default\", value=NGC_API_KEY)\n print(f\"Created secret: {NGC_SECRET_NAME}\")\nexcept ConflictError:\n print(f\"Secret '{NGC_SECRET_NAME}' already exists, continuing...\")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = \"nvidia/llama-nemotron-embed-1b-v2\"\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f\"baseline-embedding-cfg-{baseline_suffix}\"\nBASELINE_DEPLOYMENT_NAME = f\"baseline-embedding-{baseline_suffix}\"\n\nprint(\"Creating baseline deployment config...\")\nbaseline_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_CONFIG,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n image_pull_secret=NGC_SECRET_NAME,\n )\n)\n\nprint(\"Deploying base model...\")\nbaseline_deployment = client.inference.deployments.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f\"Baseline deployment: {baseline_deployment.name}\")",
+ "source": "# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get(\"NGC_API_KEY\")\nif not NGC_API_KEY:\n raise ValueError(\"NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key\")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = \"ngc-api-key\"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace=\"default\", value=NGC_API_KEY)\n print(f\"Created secret: {NGC_SECRET_NAME}\")\nexcept ConflictError:\n print(f\"Secret '{NGC_SECRET_NAME}' already exists, continuing...\")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = \"nvidia/llama-nemotron-embed-1b-v2\"\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f\"baseline-embedding-cfg-{baseline_suffix}\"\nBASELINE_DEPLOYMENT_NAME = f\"baseline-embedding-{baseline_suffix}\"\n\nprint(\"Creating baseline deployment config...\")\nbaseline_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_CONFIG,\n engine=\"nim\",\n model_spec={},\n executor_config={\n \"gpu\": 1,\n \"image_name\": NIM_IMAGE,\n \"image_tag\": NIM_TAG,\n },\n)\n\nprint(\"Deploying base model...\")\nbaseline_deployment = client.inference.deployments.create(\n workspace=\"default\",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f\"Baseline deployment: {baseline_deployment.name}\")",
"language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get("NGC_API_KEY")\nif not NGC_API_KEY:\n raise ValueError("NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = "ngc-api-key"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace="default", value=NGC_API_KEY)\n print(f"Created secret: {NGC_SECRET_NAME}")\nexcept ConflictError:\n print(f"Secret '{NGC_SECRET_NAME}' already exists, continuing...")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = "nvidia/llama-nemotron-embed-1b-v2"\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f"baseline-embedding-cfg-{baseline_suffix}"\nBASELINE_DEPLOYMENT_NAME = f"baseline-embedding-{baseline_suffix}"\n\nprint("Creating baseline deployment config...")\nbaseline_config = client.inference.deployment_configs.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_CONFIG,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n image_pull_secret=NGC_SECRET_NAME,\n )\n)\n\nprint("Deploying base model...")\nbaseline_deployment = client.inference.deployments.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f"Baseline deployment: {baseline_deployment.name}")\n"
+ "source_html": "# NGC API key is required to pull NIM images from nvcr.io\nNGC_API_KEY = os.environ.get("NGC_API_KEY")\nif not NGC_API_KEY:\n raise ValueError("NGC_API_KEY environment variable is required. Get one at https://ngc.nvidia.com/ → Setup → Generate API Key")\n\n# Create NGC secret for pulling NIM images\nNGC_SECRET_NAME = "ngc-api-key"\ntry:\n client.secrets.create(name=NGC_SECRET_NAME, workspace="default", value=NGC_API_KEY)\n print(f"Created secret: {NGC_SECRET_NAME}")\nexcept ConflictError:\n print(f"Secret '{NGC_SECRET_NAME}' already exists, continuing...")\n\n# Deploy base model for baseline comparison\nBASE_MODEL_HF = "nvidia/llama-nemotron-embed-1b-v2"\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0"\n\nbaseline_suffix = uuid.uuid4().hex[:4]\nBASELINE_DEPLOYMENT_CONFIG = f"baseline-embedding-cfg-{baseline_suffix}"\nBASELINE_DEPLOYMENT_NAME = f"baseline-embedding-{baseline_suffix}"\n\nprint("Creating baseline deployment config...")\nbaseline_config = client.inference.deployment_configs.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_CONFIG,\n engine="nim",\n model_spec={},\n executor_config={\n "gpu": 1,\n "image_name": NIM_IMAGE,\n "image_tag": NIM_TAG,\n },\n)\n\nprint("Deploying base model...")\nbaseline_deployment = client.inference.deployments.create(\n workspace="default",\n name=BASELINE_DEPLOYMENT_NAME,\n config=baseline_config.name\n)\nprint(f"Baseline deployment: {baseline_deployment.name}")\n"
},
{
"type": "code",
@@ -88,9 +90,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Create fileset to store embedding training data\nDATASET_NAME = \"embedding-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SPECTER embedding training data (scientific paper triplets)\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate upload\nprint(\"\\nUploaded files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
+ "source": "# Create fileset to store embedding training data\nDATASET_NAME = \"embedding-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SPECTER embedding training data (scientific paper triplets)\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files\nclient.files.upload(\n local_path=f\"{DATASET_PATH}/\",\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate upload\nprint(\"\\nUploaded files:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
"language": "python",
- "source_html": "# Create fileset to store embedding training data\nDATASET_NAME = "embedding-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SPECTER embedding training data (scientific paper triplets)"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate upload\nprint("\\nUploaded files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
+ "source_html": "# Create fileset to store embedding training data\nDATASET_NAME = "embedding-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SPECTER embedding training data (scientific paper triplets)"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files\nclient.files.upload(\n local_path=f"{DATASET_PATH}/",\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate upload\nprint("\\nUploaded files:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
},
{
"type": "markdown",
@@ -99,14 +101,14 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME} (created in Step 2)\")",
+ "source": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv(\"HF_TOKEN\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n workspace=\"default\",\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=\"default\")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(f\"HF_TOKEN secret: {hf_secret.name}\")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif \"NGC_SECRET_NAME\" not in globals():\n NGC_SECRET_NAME = \"ngc-api-key\"\nprint(f\"NGC_API_KEY secret: {NGC_SECRET_NAME}\")",
"language": "python",
- "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME} (created in Step 2)")\n"
+ "source_html": "# Create secrets for model access\n# Note: NGC_API_KEY secret was already created in the baseline step (Step 2)\nHF_TOKEN = os.getenv("HF_TOKEN")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n workspace="default",\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace="default")\n\n\n# Create HuggingFace token secret (for downloading model from HF during training)\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(f"HF_TOKEN secret: {hf_secret.name}")\n\n# NGC secret was already created in baseline step (Step 2), or use the platform default\nif "NGC_SECRET_NAME" not in globals():\n NGC_SECRET_NAME = "ngc-api-key"\nprint(f"NGC_API_KEY secret: {NGC_SECRET_NAME}")\n"
},
{
"type": "markdown",
- "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-3.2-nv-embedqa-1b-v2](https://huggingface.co/nvidia/llama-3.2-nv-embedqa-1b-v2) embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.\n\n---\n*Note*: Either `MODEL_NAME` or `HF_REPO_ID` below must contain the substring embed to indicate that this is an embedding model.",
- "source_html": "7. Create Base Model FileSet and Model Entity
\nCreate a fileset pointing to the nvidia/llama-3.2-nv-embedqa-1b-v2 embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n
\nNote: Either MODEL_NAME or HF_REPO_ID below must contain the substring embed to indicate that this is an embedding model.
\n"
+ "source": "### 7. Create Base Model FileSet and Model Entity\n\nCreate a fileset pointing to the [nvidia/llama-3.2-nv-embedqa-1b-v2](https://huggingface.co/nvidia/llama-3.2-nv-embedqa-1b-v2) embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.",
+ "source_html": "7. Create Base Model FileSet and Model Entity
\nCreate a fileset pointing to the nvidia/llama-3.2-nv-embedqa-1b-v2 embedding model from HuggingFace, then create a Model Entity that references this fileset. Model downloading will take place at training time.
\n"
},
{
"type": "code",
@@ -116,14 +118,14 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training_type`**: `sft` (supervised fine-tuning)\n- **Full fine-tuning**: No `peft` config needed (omit for all-weights training)\n- **`learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook creates a job with all-weights finetuning but you can also run LoRA with `merge=True`, which trains a LoRA adapter and then merges it back into the base model after training. The final output is a standard full-weight checkpoint, identical in format to an all-weights fine-tuned model, but LoRA training is faster, uses less memory, and is more lenient in hyperparameter tuning.\n\nTo do that update the job request like so\n\n```\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n peft=LoRaParamsParam(\n type= \"lora\",\n merge=True\n )\n )\n )\n)\n```",
- "source_html": "8. Create Embedding Fine-tuning Job
\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nKey hyperparameters for embedding fine-tuning:
\n\ntraining_type: sft (supervised fine-tuning) \n- Full fine-tuning: No
peft config needed (omit for all-weights training) \nlearning_rate: Lower values (1e-6 to 5e-6) work well for embedding models \nbatch_size: Larger batches improve contrastive learning (128-256 recommended) \n
\nNOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook creates a job with all-weights finetuning but you can also run LoRA with merge=True, which trains a LoRA adapter and then merges it back into the base model after training. The final output is a standard full-weight checkpoint, identical in format to an all-weights fine-tuned model, but LoRA training is faster, uses less memory, and is more lenient in hyperparameter tuning.
\nTo do that update the job request like so
\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n peft=LoRaParamsParam(\n type= "lora",\n merge=True\n )\n )\n )\n)\n
\n"
+ "source": "### 8. Create Embedding Fine-tuning Job\n\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with split `schedule`, `batch`, `optimizer`, and `parallelism` sections. Reference the model entity and dataset fileset by workspace/name (not `fileset://` URIs).\n\n**Key hyperparameters for embedding fine-tuning:**\n- **`training.training_type`**: `sft`\n- **`training.finetuning_type`**: `all_weights` for full fine-tuning, or `lora_merged` for merged LoRA\n- **`optimizer.learning_rate`**: Lower values (1e-6 to 5e-6) work well for embedding models\n- **`batch.global_batch_size`**: Larger batches improve contrastive learning (128-256 recommended)\n\n**NOTE:**\n\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set `finetuning_type` to `lora_merged`:\n\n```python\ntraining={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora_merged\",\n \"lora\": {\"rank\": 16, \"alpha\": 32},\n \"max_seq_length\": MAX_SEQ_LENGTH,\n}\n```",
+ "source_html": "8. Create Embedding Fine-tuning Job
\nCreate a customization job to fine-tune the embedding model using contrastive learning on the SPECTER dataset.
\nSubmit to the Automodel backend using AutomodelJobInput with split schedule, batch, optimizer, and parallelism sections. Reference the model entity and dataset fileset by workspace/name (not fileset:// URIs).
\nKey hyperparameters for embedding fine-tuning:
\n\ntraining.training_type: sft \ntraining.finetuning_type: all_weights for full fine-tuning, or lora_merged for merged LoRA \noptimizer.learning_rate: Lower values (1e-6 to 5e-6) work well for embedding models \nbatch.global_batch_size: Larger batches improve contrastive learning (128-256 recommended) \n
\nNOTE:
\nNeMo Platform does not support unmerged LoRA adapters for embedding models because the embedding NIM requires ONNX format, which cannot represent standalone adapters. This notebook uses all-weights fine-tuning. For merged LoRA, set finetuning_type to lora_merged:
\ntraining={\n "training_type": "sft",\n "finetuning_type": "lora_merged",\n "lora": {"rank": 16, "alpha": 32},\n "max_seq_length": MAX_SEQ_LENGTH,\n}\n
\n"
},
{
"type": "code",
- "source": "from nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"embedding-finetune-job-{job_suffix}\"\n\n# Hyperparameters optimized for embedding fine-tuning\nEPOCHS = 1\nBATCH_SIZE = 128 # Larger batches help contrastive learning\nLEARNING_RATE = 5e-6 # Lower LR for embedding models\nMAX_SEQ_LENGTH = 512 # Typical for embedding models\n\n# Note: The 'name' field must contain 'embed' for the customizer to detect this as an embedding model\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n )\n)\n\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"embedding-finetune-job-{job_suffix}\"\nOUTPUT_NAME = f\"nv-embed-finetuned-{job_suffix}\"\n\nEPOCHS = 1\nBATCH_SIZE = 128\nLEARNING_RATE = 5e-6\nMAX_SEQ_LENGTH = 512\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"all_weights\",\n \"max_seq_length\": MAX_SEQ_LENGTH,\n },\n schedule={\"epochs\": EPOCHS},\n batch={\"global_batch_size\": BATCH_SIZE, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": LEARNING_RATE},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output model: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "from nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"embedding-finetune-job-{job_suffix}"\n\n# Hyperparameters optimized for embedding fine-tuning\nEPOCHS = 1\nBATCH_SIZE = 128 # Larger batches help contrastive learning\nLEARNING_RATE = 5e-6 # Lower LR for embedding models\nMAX_SEQ_LENGTH = 512 # Typical for embedding models\n\n# Note: The 'name' field must contain 'embed' for the customizer to detect this as an embedding model\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=EPOCHS,\n batch_size=BATCH_SIZE,\n learning_rate=LEARNING_RATE,\n max_seq_length=MAX_SEQ_LENGTH,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n )\n)\n\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"embedding-finetune-job-{job_suffix}"\nOUTPUT_NAME = f"nv-embed-finetuned-{job_suffix}"\n\nEPOCHS = 1\nBATCH_SIZE = 128\nLEARNING_RATE = 5e-6\nMAX_SEQ_LENGTH = 512\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "all_weights",\n "max_seq_length": MAX_SEQ_LENGTH,\n },\n schedule={"epochs": EPOCHS},\n batch={"global_batch_size": BATCH_SIZE, "micro_batch_size": 1},\n optimizer={"learning_rate": LEARNING_RATE},\n parallelism={"num_gpus_per_node": 1},\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output model: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -132,9 +134,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
+ "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n \n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n \n time.sleep(10)",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n \n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n \n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n \n time.sleep(10)\n"
},
{
"type": "markdown",
@@ -148,15 +150,15 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace=\"default\", name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))",
+ "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace=\"default\", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))",
"language": "python",
- "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace="default", name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))\n"
+ "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace="default", name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))\n"
},
{
"type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"embedding-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"embedding-model-deployment-{deploy_suffix}\"\n\n# Embedding NIM image\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n model_name=job.spec.output.name,\n model_namespace=\"default\",\n )\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}\")",
+ "source": "# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"embedding-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"embedding-model-deployment-{deploy_suffix}\"\n\n# Embedding NIM image\nNIM_IMAGE = \"nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2\"\nNIM_TAG = \"1.13.0\" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"nim\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": OUTPUT_NAME,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": NIM_IMAGE,\n \"image_tag\": NIM_TAG,\n },\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}\")",
"language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"embedding-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"embedding-model-deployment-{deploy_suffix}"\n\n# Embedding NIM image\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=NIM_IMAGE,\n image_tag=NIM_TAG,\n gpu=1,\n model_name=job.spec.output.name,\n model_namespace="default",\n )\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}")\n"
+ "source_html": "# Create deployment config for embedding model\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"embedding-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"embedding-model-deployment-{deploy_suffix}"\n\n# Embedding NIM image\nNIM_IMAGE = "nvcr.io/nim/nvidia/llama-nemotron-embed-1b-v2"\nNIM_TAG = "1.13.0" # Update if using newer NIM release\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="nim",\n model_spec={\n "model_namespace": "default",\n "model_name": OUTPUT_NAME,\n },\n executor_config={\n "gpu": 1,\n "image_name": NIM_IMAGE,\n "image_tag": NIM_TAG,\n },\n)\n\n# Deploy model\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {client.inference.deployments.retrieve(name=deployment.name, workspace='default').status}")\n"
},
{
"type": "markdown",
@@ -176,9 +178,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f\"default/{job.spec.output.name}\"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": [DEMO_QUERY],\n \"input_type\": \"query\"\n }\n)\nquery_embedding = query_response[\"data\"][0][\"embedding\"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": DEMO_DOCS,\n \"input_type\": \"passage\"\n }\n)\ndoc_embeddings = [d[\"embedding\"] for d in doc_response[\"data\"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f\"Query: \\\"{DEMO_QUERY}\\\"\\n\")\nprint(f\"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}\")\nprint(\"-\" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f\"{DEMO_LABELS[b_idx]} [{b_score:.3f}]\" + (\" *\" if b_idx in DEMO_RELEVANT else \"\")\n f_label = f\"{DEMO_LABELS[f_idx]} [{f_score:.3f}]\" + (\" *\" if f_idx in DEMO_RELEVANT else \"\")\n \n print(f\"#{rank+1:<5} {b_label:<30} {f_label:<30}\")\n\nprint(\"\\n* = relevant paper\")\nprint(\"\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.\")",
+ "source": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f\"default/{OUTPUT_NAME}\"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": [DEMO_QUERY],\n \"input_type\": \"query\"\n }\n)\nquery_embedding = query_response[\"data\"][0][\"embedding\"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n \"v1/embeddings\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": MODEL_ID,\n \"input\": DEMO_DOCS,\n \"input_type\": \"passage\"\n }\n)\ndoc_embeddings = [d[\"embedding\"] for d in doc_response[\"data\"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f\"Query: \\\"{DEMO_QUERY}\\\"\\n\")\nprint(f\"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}\")\nprint(\"-\" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f\"{DEMO_LABELS[b_idx]} [{b_score:.3f}]\" + (\" *\" if b_idx in DEMO_RELEVANT else \"\")\n f_label = f\"{DEMO_LABELS[f_idx]} [{f_score:.3f}]\" + (\" *\" if f_idx in DEMO_RELEVANT else \"\")\n \n print(f\"#{rank+1:<5} {b_label:<30} {f_label:<30}\")\n\nprint(\"\\n* = relevant paper\")\nprint(\"\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.\")",
"language": "python",
- "source_html": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f"default/{job.spec.output.name}"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": [DEMO_QUERY],\n "input_type": "query"\n }\n)\nquery_embedding = query_response["data"][0]["embedding"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": DEMO_DOCS,\n "input_type": "passage"\n }\n)\ndoc_embeddings = [d["embedding"] for d in doc_response["data"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f"Query: \\"{DEMO_QUERY}\\"\\n")\nprint(f"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}")\nprint("-" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f"{DEMO_LABELS[b_idx]} [{b_score:.3f}]" + (" *" if b_idx in DEMO_RELEVANT else "")\n f_label = f"{DEMO_LABELS[f_idx]} [{f_score:.3f}]" + (" *" if f_idx in DEMO_RELEVANT else "")\n \n print(f"#{rank+1:<5} {b_label:<30} {f_label:<30}")\n\nprint("\\n* = relevant paper")\nprint("\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.")\n"
+ "source_html": "# Compare: same query, base model vs fine-tuned\n# Using the same DEMO_QUERY and DEMO_DOCS from the baseline test\nMODEL_ID = f"default/{OUTPUT_NAME}"\n\n# Get query embedding from fine-tuned model\nquery_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": [DEMO_QUERY],\n "input_type": "query"\n }\n)\nquery_embedding = query_response["data"][0]["embedding"]\n\n# Get document embeddings from fine-tuned model\ndoc_response = client.inference.gateway.provider.post(\n "v1/embeddings",\n name=deployment.name,\n workspace="default",\n body={\n "model": MODEL_ID,\n "input": DEMO_DOCS,\n "input_type": "passage"\n }\n)\ndoc_embeddings = [d["embedding"] for d in doc_response["data"]]\n\n# Calculate similarities and rank\nscores = [(i, cosine_similarity(query_embedding, doc_embeddings[i])) for i in range(len(DEMO_DOCS))]\nFINETUNED_RANKING = sorted(scores, key=lambda x: -x[1])\n\n# Display side-by-side comparison\nprint(f"Query: \\"{DEMO_QUERY}\\"\\n")\nprint(f"{'Rank':<6} {'Base Model':<30} {'Fine-tuned Model':<30}")\nprint("-" * 66)\n\nfor rank in range(len(DEMO_DOCS)):\n b_idx, b_score = BASELINE_RANKING[rank]\n f_idx, f_score = FINETUNED_RANKING[rank]\n \n b_label = f"{DEMO_LABELS[b_idx]} [{b_score:.3f}]" + (" *" if b_idx in DEMO_RELEVANT else "")\n f_label = f"{DEMO_LABELS[f_idx]} [{f_score:.3f}]" + (" *" if f_idx in DEMO_RELEVANT else "")\n \n print(f"#{rank+1:<5} {b_label:<30} {f_label:<30}")\n\nprint("\\n* = relevant paper")\nprint("\\nThe fine-tuned model pushes 'Random Forest' down and ranks CRF papers higher.")\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/lora-customization-job.json b/docs/fern/components/notebooks/lora-customization-job.json
index bb1ac4e94d..e2f2ff8b16 100644
--- a/docs/fern/components/notebooks/lora-customization-job.json
+++ b/docs/fern/components/notebooks/lora-customization-job.json
@@ -7,8 +7,8 @@
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)\n3. **Installed the `datasets` package** for loading SQuAD: `pip install datasets`\n4. **At least one GPU with CUDA 12.8+**",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n- Installed the
datasets package for loading SQuAD: pip install datasets \n- At least one GPU with CUDA 12.8+
\n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **Installed the `datasets` package** for loading SQuAD: `pip install datasets`\n4. **At least one GPU with CUDA 12.8+**",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n- Installed the
datasets package for loading SQuAD: pip install datasets \n- At least one GPU with CUDA 12.8+
\n
\n"
},
{
"type": "markdown",
@@ -17,9 +17,9 @@
},
{
"type": "code",
- "source": "import os\nimport json\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n DeploymentParamsParam,\n LoRaParamsParam,\n ParallelismParamsParam,\n SftTrainingParam,\n)\n\n\ndef sanitize_name(prefix: str, name: str):\n \"\"\"Sanitize model_name for deployment/config naming. Compatible with platform naming rules.\"\"\"\n name = name.split(\"/\")[-1]\n sanitized = re.sub(r\"[^a-z0-9@.+_-]\", \"-\", name.lower())\n sanitized = re.sub(r\"-+\", \"-\", sanitized).strip(\"-\")\n return f\"{prefix}-{sanitized}\"[:59].rstrip(\"-\")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = \"\"):\n \"\"\"Return a check() that raises TimeoutError if called after `seconds` have elapsed.\"\"\"\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f\"{job_name} took longer than {seconds} seconds\")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace=\"default\"\n)",
+ "source": "import json\nimport os\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\n\ndef sanitize_name(prefix: str, name: str):\n \"\"\"Sanitize model_name for deployment/config naming. Compatible with platform naming rules.\"\"\"\n name = name.split(\"/\")[-1]\n sanitized = re.sub(r\"[^a-z0-9@.+_-]\", \"-\", name.lower())\n sanitized = re.sub(r\"-+\", \"-\", sanitized).strip(\"-\")\n return f\"{prefix}-{sanitized}\"[:59].rstrip(\"-\")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = \"\"):\n \"\"\"Return a check() that raises TimeoutError if called after `seconds` have elapsed.\"\"\"\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f\"{job_name} took longer than {seconds} seconds\")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace=\"default\"\n)",
"language": "python",
- "source_html": "import os\nimport json\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n DeploymentParamsParam,\n LoRaParamsParam,\n ParallelismParamsParam,\n SftTrainingParam,\n)\n\n\ndef sanitize_name(prefix: str, name: str):\n """Sanitize model_name for deployment/config naming. Compatible with platform naming rules."""\n name = name.split("/")[-1]\n sanitized = re.sub(r"[^a-z0-9@.+_-]", "-", name.lower())\n sanitized = re.sub(r"-+", "-", sanitized).strip("-")\n return f"{prefix}-{sanitized}"[:59].rstrip("-")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = ""):\n """Return a check() that raises TimeoutError if called after `seconds` have elapsed."""\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f"{job_name} took longer than {seconds} seconds")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace="default"\n)\n"
+ "source_html": "import json\nimport os\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\n\ndef sanitize_name(prefix: str, name: str):\n """Sanitize model_name for deployment/config naming. Compatible with platform naming rules."""\n name = name.split("/")[-1]\n sanitized = re.sub(r"[^a-z0-9@.+_-]", "-", name.lower())\n sanitized = re.sub(r"-+", "-", sanitized).strip("-")\n return f"{prefix}-{sanitized}"[:59].rstrip("-")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = ""):\n """Return a check() that raises TimeoutError if called after `seconds` have elapsed."""\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f"{job_name} took longer than {seconds} seconds")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace="default"\n)\n"
},
{
"type": "markdown",
@@ -77,14 +77,14 @@
},
{
"type": "markdown",
- "source": "### 6. Create LoRA Customization Job\n\nSubmit a customization job with `training=SftTrainingParam(type=\"sft\", peft=LoRaParamsParam(type=\"lora\"), ...)`. Set `lora_enabled=True` in the `deployment_config` so the platform can deploy the base model with LoRA support automatically.\n\nWhen LoRA Enabled is set to true for Models Deployed via the `/apis/models` endpoint or via the `deployment_config` option during the customization job, all LoRA adapters (enabled by default) will get automatically deployed in the NIM.",
- "source_html": "6. Create LoRA Customization Job
\nSubmit a customization job with training=SftTrainingParam(type="sft", peft=LoRaParamsParam(type="lora"), ...). Set lora_enabled=True in the deployment_config so the platform can deploy the base model with LoRA support automatically.
\nWhen LoRA Enabled is set to true for Models Deployed via the /apis/models endpoint or via the deployment_config option during the customization job, all LoRA adapters (enabled by default) will get automatically deployed in the NIM.
\n"
+ "source": "### 6. Create LoRA Customization Job\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with `finetuning_type: lora`. After training completes, deploy the base model with LoRA support manually (step 8).",
+ "source_html": "6. Create LoRA Customization Job
\nSubmit to the Automodel backend using AutomodelJobInput with finetuning_type: lora. After training completes, deploy the base model with LoRA support manually (step 8).
\n"
},
{
"type": "code",
- "source": "job_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n context_parallel_size=1,\n expert_parallel_size=1,\n ),\n micro_batch_size=1,\n peft=LoRaParamsParam(type=\"lora\"),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n gpu=1,\n additional_envs={\"NIM_MODEL_PROFILE\": \"vllm-lora\"},\n ),\n ),\n)\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
+ "source": "from nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\nOUTPUT_NAME = f\"lora-adapter-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 2048,\n },\n schedule={\"epochs\": 2},\n batch={\"global_batch_size\": 64, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\n \"num_gpus_per_node\": 1,\n \"num_nodes\": 1,\n \"tensor_parallel_size\": 1,\n \"pipeline_parallel_size\": 1,\n \"context_parallel_size\": 1,\n \"expert_parallel_size\": 1,\n },\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output adapter: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "job_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"my-sft-job-{job_suffix}"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n context_parallel_size=1,\n expert_parallel_size=1,\n ),\n micro_batch_size=1,\n peft=LoRaParamsParam(type="lora"),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n gpu=1,\n additional_envs={"NIM_MODEL_PROFILE": "vllm-lora"},\n ),\n ),\n)\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
+ "source_html": "from nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"my-sft-job-{job_suffix}"\nOUTPUT_NAME = f"lora-adapter-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 2048,\n },\n schedule={"epochs": 2},\n batch={"global_batch_size": 64, "micro_batch_size": 1},\n optimizer={"learning_rate": 5e-5},\n parallelism={\n "num_gpus_per_node": 1,\n "num_nodes": 1,\n "tensor_parallel_size": 1,\n "pipeline_parallel_size": 1,\n "context_parallel_size": 1,\n "expert_parallel_size": 1,\n },\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output adapter: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -93,20 +93,20 @@
},
{
"type": "code",
- "source": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, \"Customization Job\")\nwhile True:\n time_check()\n status = client.customization.jobs.get_status(name=job.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Job Status: {status.status}\")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get(\"step\"), d.get(\"max_steps\")\n training_phase = d.get(\"phase\")\n break\n break\n if step is not None and max_steps is not None:\n print(f\"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)\")\n if training_phase:\n print(f\"Phase: {training_phase}\")\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished: {status.status}\")\n break\n time.sleep(10)\n\nassert status.status == \"completed\"",
+ "source": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, \"Customization Job\")\nwhile True:\n time_check()\n status = client.jobs.get_status(name=job.job.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Job Status: {status.status}\")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get(\"step\"), d.get(\"max_steps\")\n training_phase = d.get(\"phase\")\n break\n break\n if step is not None and max_steps is not None:\n print(f\"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)\")\n if training_phase:\n print(f\"Phase: {training_phase}\")\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished: {status.status}\")\n break\n time.sleep(10)\n\nassert status.status == \"completed\"",
"language": "python",
- "source_html": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, "Customization Job")\nwhile True:\n time_check()\n status = client.customization.jobs.get_status(name=job.name, workspace="default")\n clear_output(wait=True)\n print(f"Job Status: {status.status}")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get("step"), d.get("max_steps")\n training_phase = d.get("phase")\n break\n break\n if step is not None and max_steps is not None:\n print(f"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)")\n if training_phase:\n print(f"Phase: {training_phase}")\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished: {status.status}")\n break\n time.sleep(10)\n\nassert status.status == "completed"\n"
+ "source_html": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, "Customization Job")\nwhile True:\n time_check()\n status = client.jobs.get_status(name=job.job.name, workspace="default")\n clear_output(wait=True)\n print(f"Job Status: {status.status}")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get("step"), d.get("max_steps")\n training_phase = d.get("phase")\n break\n break\n if step is not None and max_steps is not None:\n print(f"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)")\n if training_phase:\n print(f"Phase: {training_phase}")\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished: {status.status}")\n break\n time.sleep(10)\n\nassert status.status == "completed"\n"
},
{
"type": "markdown",
- "source": "### 8. Validate Output Model and Deployment\n\nWith `deployment_config` configured, the platform will create a NIM deployment for the base model after training. The fine-tuned LoRA adapter is enabled by default and automatically served through the deployment. Check the model entity and deployment status.",
- "source_html": "8. Validate Output Model and Deployment
\nWith deployment_config configured, the platform will create a NIM deployment for the base model after training. The fine-tuned LoRA adapter is enabled by default and automatically served through the deployment. Check the model entity and deployment status.
\n"
+ "source": "### 8. Validate Output Model and Deployment\n\nWith the base model entity and LoRA adapter from training, create a NIM deployment with `lora_enabled=True` so the adapter is served alongside the base weights. Check the model entity and deployment status.",
+ "source_html": "8. Validate Output Model and Deployment
\nWith the base model entity and LoRA adapter from training, create a NIM deployment with lora_enabled=True so the adapter is served alongside the base weights. Check the model entity and deployment status.
\n"
},
{
"type": "code",
- "source": "model_entity = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n# Clear verbose linear_layers list for cleaner output\nmodel_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\n\ndeployment_name = sanitize_name(\"sft-deploy\", job.spec.model)\ndeployment_status = client.inference.deployments.retrieve(name=deployment_name, workspace=\"default\")\nprint(f\"Deployment status: {deployment_status.status}\")",
+ "source": "deploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"lora-deploy-cfg-{deploy_suffix}\"\ndeployment_name = f\"lora-deploy-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": MODEL_NAME,\n \"lora_enabled\": True,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n \"additional_args\": [\"--max-lora-rank\", \"32\"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=deployment_name,\n config=deployment_config.name,\n)\n\nmodel_entity = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\nif model_entity.spec:\n model_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\nprint(f\"Deployment status: {deployment.status}\")",
"language": "python",
- "source_html": "model_entity = client.models.retrieve(workspace="default", name=MODEL_NAME)\n# Clear verbose linear_layers list for cleaner output\nmodel_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\n\ndeployment_name = sanitize_name("sft-deploy", job.spec.model)\ndeployment_status = client.inference.deployments.retrieve(name=deployment_name, workspace="default")\nprint(f"Deployment status: {deployment_status.status}")\n"
+ "source_html": "deploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"lora-deploy-cfg-{deploy_suffix}"\ndeployment_name = f"lora-deploy-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={\n "model_namespace": "default",\n "model_name": MODEL_NAME,\n "lora_enabled": True,\n },\n executor_config={\n "gpu": 1,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n "additional_args": ["--max-lora-rank", "32"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=deployment_name,\n config=deployment_config.name,\n)\n\nmodel_entity = client.models.retrieve(workspace="default", name=MODEL_NAME)\nif model_entity.spec:\n model_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\nprint(f"Deployment status: {deployment.status}")\n"
},
{
"type": "markdown",
@@ -126,14 +126,14 @@
},
{
"type": "code",
- "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": job.spec.output.name,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
+ "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": OUTPUT_NAME,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
"language": "python",
- "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": job.spec.output.name,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
+ "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": OUTPUT_NAME,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
},
{
"type": "markdown",
- "source": "## Conclusion\n\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the `output.name` to access the model for further inference or evaluation.\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Try [Full SFT](./sft-customization-job) or [DPO](./dpo-customization-job) for other customization options",
- "source_html": "Conclusion
\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the output.name to access the model for further inference or evaluation.
\nNext Steps
\n\n"
+ "source": "## Conclusion\n\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the `output.name` to access the model for further inference or evaluation.\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Try [Full SFT](./sft-customization-job) for other customization options",
+ "source_html": "Conclusion
\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the output.name to access the model for further inference or evaluation.
\nNext Steps
\n\n"
}
]
}
\ No newline at end of file
diff --git a/docs/fern/components/notebooks/lora-customization-job.ts b/docs/fern/components/notebooks/lora-customization-job.ts
index 3f1aacbf4e..d4b0272cd4 100644
--- a/docs/fern/components/notebooks/lora-customization-job.ts
+++ b/docs/fern/components/notebooks/lora-customization-job.ts
@@ -1,7 +1,9 @@
-// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
-// SPDX-License-Identifier: Apache-2.0
-
-/** Auto-generated by ipynb-to-fern-json.py - do not edit */
+/**
+ * SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+ * SPDX-License-Identifier: Apache-2.0
+ *
+ * Auto-generated by ipynb-to-fern-json.py - do not edit manually.
+ */
export default { cells: [
{
"type": "markdown",
@@ -10,8 +12,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)\n3. **Installed the `datasets` package** for loading SQuAD: `pip install datasets`\n4. **At least one GPU with CUDA 12.8+**",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n- Installed the
datasets package for loading SQuAD: pip install datasets \n- At least one GPU with CUDA 12.8+
\n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n3. **Installed the `datasets` package** for loading SQuAD: `pip install datasets`\n4. **At least one GPU with CUDA 12.8+**",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n- Installed the
datasets package for loading SQuAD: pip install datasets \n- At least one GPU with CUDA 12.8+
\n
\n"
},
{
"type": "markdown",
@@ -20,9 +22,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "import os\nimport json\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n DeploymentParamsParam,\n LoRaParamsParam,\n ParallelismParamsParam,\n SftTrainingParam,\n)\n\n\ndef sanitize_name(prefix: str, name: str):\n \"\"\"Sanitize model_name for deployment/config naming. Compatible with platform naming rules.\"\"\"\n name = name.split(\"/\")[-1]\n sanitized = re.sub(r\"[^a-z0-9@.+_-]\", \"-\", name.lower())\n sanitized = re.sub(r\"-+\", \"-\", sanitized).strip(\"-\")\n return f\"{prefix}-{sanitized}\"[:59].rstrip(\"-\")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = \"\"):\n \"\"\"Return a check() that raises TimeoutError if called after `seconds` have elapsed.\"\"\"\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f\"{job_name} took longer than {seconds} seconds\")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace=\"default\"\n)",
+ "source": "import json\nimport os\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\n\ndef sanitize_name(prefix: str, name: str):\n \"\"\"Sanitize model_name for deployment/config naming. Compatible with platform naming rules.\"\"\"\n name = name.split(\"/\")[-1]\n sanitized = re.sub(r\"[^a-z0-9@.+_-]\", \"-\", name.lower())\n sanitized = re.sub(r\"-+\", \"-\", sanitized).strip(\"-\")\n return f\"{prefix}-{sanitized}\"[:59].rstrip(\"-\")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = \"\"):\n \"\"\"Return a check() that raises TimeoutError if called after `seconds` have elapsed.\"\"\"\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f\"{job_name} took longer than {seconds} seconds\")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get(\"NMP_BASE_URL\", \"http://localhost:8080\")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace=\"default\"\n)",
"language": "python",
- "source_html": "import os\nimport json\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n DeploymentParamsParam,\n LoRaParamsParam,\n ParallelismParamsParam,\n SftTrainingParam,\n)\n\n\ndef sanitize_name(prefix: str, name: str):\n """Sanitize model_name for deployment/config naming. Compatible with platform naming rules."""\n name = name.split("/")[-1]\n sanitized = re.sub(r"[^a-z0-9@.+_-]", "-", name.lower())\n sanitized = re.sub(r"-+", "-", sanitized).strip("-")\n return f"{prefix}-{sanitized}"[:59].rstrip("-")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = ""):\n """Return a check() that raises TimeoutError if called after `seconds` have elapsed."""\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f"{job_name} took longer than {seconds} seconds")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace="default"\n)\n"
+ "source_html": "import json\nimport os\nimport re\nimport time\nimport uuid\nfrom pathlib import Path\nfrom nemo_platform import NeMoPlatform, ConflictError\nfrom nemo_platform.types.secrets import PlatformSecretResponse\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\n\ndef sanitize_name(prefix: str, name: str):\n """Sanitize model_name for deployment/config naming. Compatible with platform naming rules."""\n name = name.split("/")[-1]\n sanitized = re.sub(r"[^a-z0-9@.+_-]", "-", name.lower())\n sanitized = re.sub(r"-+", "-", sanitized).strip("-")\n return f"{prefix}-{sanitized}"[:59].rstrip("-")\n\n\ndef max_wait_time_checker(seconds: int, job_name: str = ""):\n """Return a check() that raises TimeoutError if called after `seconds` have elapsed."""\n start_time = time.time()\n\n def check():\n if time.time() - start_time > seconds:\n raise TimeoutError(f"{job_name} took longer than {seconds} seconds")\n\n return check\n\n\nNMP_BASE_URL = os.environ.get("NMP_BASE_URL", "http://localhost:8080")\nclient = NeMoPlatform(\n base_url=NMP_BASE_URL,\n workspace="default"\n)\n"
},
{
"type": "markdown",
@@ -80,14 +82,14 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "### 6. Create LoRA Customization Job\n\nSubmit a customization job with `training=SftTrainingParam(type=\"sft\", peft=LoRaParamsParam(type=\"lora\"), ...)`. Set `lora_enabled=True` in the `deployment_config` so the platform can deploy the base model with LoRA support automatically.\n\nWhen LoRA Enabled is set to true for Models Deployed via the `/apis/models` endpoint or via the `deployment_config` option during the customization job, all LoRA adapters (enabled by default) will get automatically deployed in the NIM.",
- "source_html": "6. Create LoRA Customization Job
\nSubmit a customization job with training=SftTrainingParam(type="sft", peft=LoRaParamsParam(type="lora"), ...). Set lora_enabled=True in the deployment_config so the platform can deploy the base model with LoRA support automatically.
\nWhen LoRA Enabled is set to true for Models Deployed via the /apis/models endpoint or via the deployment_config option during the customization job, all LoRA adapters (enabled by default) will get automatically deployed in the NIM.
\n"
+ "source": "### 6. Create LoRA Customization Job\n\nSubmit to the **Automodel** backend using `AutomodelJobInput` with `finetuning_type: lora`. After training completes, deploy the base model with LoRA support manually (step 8).",
+ "source_html": "6. Create LoRA Customization Job
\nSubmit to the Automodel backend using AutomodelJobInput with finetuning_type: lora. After training completes, deploy the base model with LoRA support manually (step 8).
\n"
},
{
"type": "code",
- "source": "job_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n context_parallel_size=1,\n expert_parallel_size=1,\n ),\n micro_batch_size=1,\n peft=LoRaParamsParam(type=\"lora\"),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n gpu=1,\n additional_envs={\"NIM_MODEL_PROFILE\": \"vllm-lora\"},\n ),\n ),\n)\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
+ "source": "from nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\nOUTPUT_NAME = f\"lora-adapter-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 2048,\n },\n schedule={\"epochs\": 2},\n batch={\"global_batch_size\": 64, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\n \"num_gpus_per_node\": 1,\n \"num_nodes\": 1,\n \"tensor_parallel_size\": 1,\n \"pipeline_parallel_size\": 1,\n \"context_parallel_size\": 1,\n \"expert_parallel_size\": 1,\n },\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output adapter: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "job_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"my-sft-job-{job_suffix}"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n context_parallel_size=1,\n expert_parallel_size=1,\n ),\n micro_batch_size=1,\n peft=LoRaParamsParam(type="lora"),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n gpu=1,\n additional_envs={"NIM_MODEL_PROFILE": "vllm-lora"},\n ),\n ),\n)\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
+ "source_html": "from nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"my-sft-job-{job_suffix}"\nOUTPUT_NAME = f"lora-adapter-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 2048,\n },\n schedule={"epochs": 2},\n batch={"global_batch_size": 64, "micro_batch_size": 1},\n optimizer={"learning_rate": 5e-5},\n parallelism={\n "num_gpus_per_node": 1,\n "num_nodes": 1,\n "tensor_parallel_size": 1,\n "pipeline_parallel_size": 1,\n "context_parallel_size": 1,\n "expert_parallel_size": 1,\n },\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output adapter: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -96,20 +98,20 @@ export default { cells: [
},
{
"type": "code",
- "source": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, \"Customization Job\")\nwhile True:\n time_check()\n status = client.customization.jobs.get_status(name=job.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Job Status: {status.status}\")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get(\"step\"), d.get(\"max_steps\")\n training_phase = d.get(\"phase\")\n break\n break\n if step is not None and max_steps is not None:\n print(f\"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)\")\n if training_phase:\n print(f\"Phase: {training_phase}\")\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished: {status.status}\")\n break\n time.sleep(10)\n\nassert status.status == \"completed\"",
+ "source": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, \"Customization Job\")\nwhile True:\n time_check()\n status = client.jobs.get_status(name=job.job.name, workspace=\"default\")\n clear_output(wait=True)\n print(f\"Job Status: {status.status}\")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get(\"step\"), d.get(\"max_steps\")\n training_phase = d.get(\"phase\")\n break\n break\n if step is not None and max_steps is not None:\n print(f\"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)\")\n if training_phase:\n print(f\"Phase: {training_phase}\")\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished: {status.status}\")\n break\n time.sleep(10)\n\nassert status.status == \"completed\"",
"language": "python",
- "source_html": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, "Customization Job")\nwhile True:\n time_check()\n status = client.customization.jobs.get_status(name=job.name, workspace="default")\n clear_output(wait=True)\n print(f"Job Status: {status.status}")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get("step"), d.get("max_steps")\n training_phase = d.get("phase")\n break\n break\n if step is not None and max_steps is not None:\n print(f"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)")\n if training_phase:\n print(f"Phase: {training_phase}")\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished: {status.status}")\n break\n time.sleep(10)\n\nassert status.status == "completed"\n"
+ "source_html": "from IPython.display import clear_output\n\ntime_check = max_wait_time_checker(3600, "Customization Job")\nwhile True:\n time_check()\n status = client.jobs.get_status(name=job.job.name, workspace="default")\n clear_output(wait=True)\n print(f"Job Status: {status.status}")\n step = max_steps = training_phase = None\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n d = task.status_details or {}\n step, max_steps = d.get("step"), d.get("max_steps")\n training_phase = d.get("phase")\n break\n break\n if step is not None and max_steps is not None:\n print(f"Training: Step {step}/{max_steps} ({100 * step / max_steps:.1f}%)")\n if training_phase:\n print(f"Phase: {training_phase}")\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished: {status.status}")\n break\n time.sleep(10)\n\nassert status.status == "completed"\n"
},
{
"type": "markdown",
- "source": "### 8. Validate Output Model and Deployment\n\nWith `deployment_config` configured, the platform will create a NIM deployment for the base model after training. The fine-tuned LoRA adapter is enabled by default and automatically served through the deployment. Check the model entity and deployment status.",
- "source_html": "8. Validate Output Model and Deployment
\nWith deployment_config configured, the platform will create a NIM deployment for the base model after training. The fine-tuned LoRA adapter is enabled by default and automatically served through the deployment. Check the model entity and deployment status.
\n"
+ "source": "### 8. Validate Output Model and Deployment\n\nWith the base model entity and LoRA adapter from training, create a NIM deployment with `lora_enabled=True` so the adapter is served alongside the base weights. Check the model entity and deployment status.",
+ "source_html": "8. Validate Output Model and Deployment
\nWith the base model entity and LoRA adapter from training, create a NIM deployment with lora_enabled=True so the adapter is served alongside the base weights. Check the model entity and deployment status.
\n"
},
{
"type": "code",
- "source": "model_entity = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\n# Clear verbose linear_layers list for cleaner output\nmodel_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\n\ndeployment_name = sanitize_name(\"sft-deploy\", job.spec.model)\ndeployment_status = client.inference.deployments.retrieve(name=deployment_name, workspace=\"default\")\nprint(f\"Deployment status: {deployment_status.status}\")",
+ "source": "deploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"lora-deploy-cfg-{deploy_suffix}\"\ndeployment_name = f\"lora-deploy-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": MODEL_NAME,\n \"lora_enabled\": True,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n \"additional_args\": [\"--max-lora-rank\", \"32\"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=deployment_name,\n config=deployment_config.name,\n)\n\nmodel_entity = client.models.retrieve(workspace=\"default\", name=MODEL_NAME)\nif model_entity.spec:\n model_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\nprint(f\"Deployment status: {deployment.status}\")",
"language": "python",
- "source_html": "model_entity = client.models.retrieve(workspace="default", name=MODEL_NAME)\n# Clear verbose linear_layers list for cleaner output\nmodel_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\n\ndeployment_name = sanitize_name("sft-deploy", job.spec.model)\ndeployment_status = client.inference.deployments.retrieve(name=deployment_name, workspace="default")\nprint(f"Deployment status: {deployment_status.status}")\n"
+ "source_html": "deploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"lora-deploy-cfg-{deploy_suffix}"\ndeployment_name = f"lora-deploy-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={\n "model_namespace": "default",\n "model_name": MODEL_NAME,\n "lora_enabled": True,\n },\n executor_config={\n "gpu": 1,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n "additional_args": ["--max-lora-rank", "32"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=deployment_name,\n config=deployment_config.name,\n)\n\nmodel_entity = client.models.retrieve(workspace="default", name=MODEL_NAME)\nif model_entity.spec:\n model_entity.spec.linear_layers = None\nprint(model_entity.model_dump_json(indent=2))\nprint(f"Deployment status: {deployment.status}")\n"
},
{
"type": "markdown",
@@ -129,13 +131,13 @@ export default { cells: [
},
{
"type": "code",
- "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": job.spec.output.name,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
+ "source": "context = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=\"default\",\n body={\n \"model\": OUTPUT_NAME,\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 256,\n }\n)\nprint(\"=\" * 60)\nprint(\"MODEL INFERENCE\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
"language": "python",
- "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": job.spec.output.name,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
+ "source_html": "context = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace="default",\n body={\n "model": OUTPUT_NAME,\n "messages": messages,\n "temperature": 0,\n "max_tokens": 256,\n }\n)\nprint("=" * 60)\nprint("MODEL INFERENCE")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
},
{
"type": "markdown",
- "source": "## Conclusion\n\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the `output.name` to access the model for further inference or evaluation.\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Try [Full SFT](./sft-customization-job) or [DPO](./dpo-customization-job) for other customization options",
- "source_html": "Conclusion
\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the output.name to access the model for further inference or evaluation.
\nNext Steps
\n\n"
+ "source": "## Conclusion\n\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the `output.name` to access the model for further inference or evaluation.\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Try [Full SFT](./sft-customization-job) for other customization options",
+ "source_html": "Conclusion
\nYou have started a LoRA customization job, monitored it to completion, and evaluated the fine-tuned model. Use the output.name to access the model for further inference or evaluation.
\nNext Steps
\n\n"
}
] };
diff --git a/docs/fern/components/notebooks/optimize-throughput.json b/docs/fern/components/notebooks/optimize-throughput.json
index 17d520aa81..452ccf038c 100644
--- a/docs/fern/components/notebooks/optimize-throughput.json
+++ b/docs/fern/components/notebooks/optimize-throughput.json
@@ -7,8 +7,8 @@
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n
\n"
},
{
"type": "markdown",
@@ -50,9 +50,9 @@
},
{
"type": "code",
- "source": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
+ "source": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f\"{DATASET_PATH}/\", # Trailing slash uploads directory contents to fileset root\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
"language": "python",
- "source_html": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
+ "source_html": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f"{DATASET_PATH}/", # Trailing slash uploads directory contents to fileset root\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
},
{
"type": "markdown",
@@ -78,14 +78,14 @@
},
{
"type": "markdown",
- "source": "### 5. Create LoRA Job with Sequence Packing\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.",
- "source_html": "5. Create LoRA Job with Sequence Packing
\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.
\n"
+ "source": "### 5. Create LoRA Job with Sequence Packing\nCreate a LoRA customization job with **sequence packing** enabled via `AutomodelJobInput` (`batch.sequence_packing=True`).",
+ "source_html": "5. Create LoRA Job with Sequence Packing
\nCreate a LoRA customization job with sequence packing enabled via AutomodelJobInput (batch.sequence_packing=True).
\n"
},
{
"type": "code",
- "source": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n)\n\n# Enable sequence packing to improve throughput and GPU utilization\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\njob_spec = CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=SEQUENCE_PACKING_ENABLED,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_with_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=job_spec\n)\n\nprint(f\"Job ID: {job_with_sequence_packing.name}\")\nprint(f\"Output model: {job_with_sequence_packing.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"packing-job-{job_suffix}\"\nPACK_OUTPUT_NAME = f\"packing-out-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 4096,\n },\n schedule={\"epochs\": 1, \"val_check_interval\": 0.1},\n batch={\n \"global_batch_size\": 64,\n \"micro_batch_size\": 1,\n \"sequence_packing\": SEQUENCE_PACKING_ENABLED,\n },\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": PACK_OUTPUT_NAME},\n)\n\njob_with_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job_with_sequence_packing.job.name}\")\nprint(f\"Output adapter: {PACK_OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n)\n\n# Enable sequence packing to improve throughput and GPU utilization\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-sft-job-{job_suffix}"\njob_spec = CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=SEQUENCE_PACKING_ENABLED,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_with_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=job_spec\n)\n\nprint(f"Job ID: {job_with_sequence_packing.name}")\nprint(f"Output model: {job_with_sequence_packing.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"packing-job-{job_suffix}"\nPACK_OUTPUT_NAME = f"packing-out-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 4096,\n },\n schedule={"epochs": 1, "val_check_interval": 0.1},\n batch={\n "global_batch_size": 64,\n "micro_batch_size": 1,\n "sequence_packing": SEQUENCE_PACKING_ENABLED,\n },\n optimizer={"learning_rate": 5e-5},\n parallelism={"num_gpus_per_node": 1},\n output={"name": PACK_OUTPUT_NAME},\n)\n\njob_with_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job_with_sequence_packing.job.name}")\nprint(f"Output adapter: {PACK_OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -110,20 +110,20 @@
},
{
"type": "code",
- "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"loss\"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n val_loss = float(td[val_loss_key]) if val_loss_key in td else None\n train_loss = float(td[train_loss_key]) if train_loss_key in td else None\n break\n break\n\n # Fall back to top-level status_details\n if status.status_details:\n if val_loss is None and val_loss_key in status.status_details:\n val_loss = float(status.status_details[val_loss_key])\n if train_loss is None and train_loss_key in status.status_details:\n train_loss = float(status.status_details[train_loss_key])\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.name,\n timeout=TIMEOUT_SECONDS,\n)\n\nprint(f\"Validation loss: {job_with_sequence_packing_status.status_details['val_loss']:.2f}\")",
+ "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"loss\"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get(\"val_loss\")\nif packed_val_loss is not None:\n print(f\"Validation loss: {float(packed_val_loss):.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")",
"language": "python",
- "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "loss"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n val_loss = float(td[val_loss_key]) if val_loss_key in td else None\n train_loss = float(td[train_loss_key]) if train_loss_key in td else None\n break\n break\n\n # Fall back to top-level status_details\n if status.status_details:\n if val_loss is None and val_loss_key in status.status_details:\n val_loss = float(status.status_details[val_loss_key])\n if train_loss is None and train_loss_key in status.status_details:\n train_loss = float(status.status_details[train_loss_key])\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.name,\n timeout=TIMEOUT_SECONDS,\n)\n\nprint(f"Validation loss: {job_with_sequence_packing_status.status_details['val_loss']:.2f}")\n"
+ "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "loss"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "training":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get("val_loss")\nif packed_val_loss is not None:\n print(f"Validation loss: {float(packed_val_loss):.2f}")\nelse:\n print("Validation loss: not reported in job status")\n"
},
{
"type": "markdown",
- "source": "### 7. Create LoRA Job without Sequence Packing\nCreate a customization job with sequence packing disabled. It's expected to take longer to complete.",
- "source_html": "7. Create LoRA Job without Sequence Packing
\nCreate a customization job with sequence packing disabled. It's expected to take longer to complete.
\n"
+ "source": "### 7. Create LoRA Job without Sequence Packing\nCreate a second Automodel LoRA job with `batch.sequence_packing=False` for comparison.",
+ "source_html": "7. Create LoRA Job without Sequence Packing
\nCreate a second Automodel LoRA job with batch.sequence_packing=False for comparison.
\n"
},
{
"type": "code",
- "source": "import uuid\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\n\njob_spec_no_packing = CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=False,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_without_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=job_spec_no_packing\n)\n\nprint(f\"Job ID: {job_without_sequence_packing.name}\")\nprint(f\"Output model: {job_without_sequence_packing.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"no-packing-job-{job_suffix}\"\nNO_PACK_OUTPUT_NAME = f\"no-packing-out-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 4096,\n },\n schedule={\"epochs\": 1, \"val_check_interval\": 0.1},\n batch={\n \"global_batch_size\": 64,\n \"micro_batch_size\": 1,\n \"sequence_packing\": False,\n },\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": NO_PACK_OUTPUT_NAME},\n)\n\njob_without_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job_without_sequence_packing.job.name}\")\nprint(f\"Output adapter: {NO_PACK_OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"my-sft-job-{job_suffix}"\n\njob_spec_no_packing = CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=False,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_without_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=job_spec_no_packing\n)\n\nprint(f"Job ID: {job_without_sequence_packing.name}")\nprint(f"Output model: {job_without_sequence_packing.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"no-packing-job-{job_suffix}"\nNO_PACK_OUTPUT_NAME = f"no-packing-out-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 4096,\n },\n schedule={"epochs": 1, "val_check_interval": 0.1},\n batch={\n "global_batch_size": 64,\n "micro_batch_size": 1,\n "sequence_packing": False,\n },\n optimizer={"learning_rate": 5e-5},\n parallelism={"num_gpus_per_node": 1},\n output={"name": NO_PACK_OUTPUT_NAME},\n)\n\njob_without_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job_without_sequence_packing.job.name}")\nprint(f"Output adapter: {NO_PACK_OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -132,9 +132,9 @@
},
{
"type": "code",
- "source": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_without_sequence_packing.name,\n timeout=TIMEOUT_SECONDS\n)\n\nprint(f\"Validation loss: {job_without_sequence_packing_status.status_details['val_loss']:.2f}\")",
+ "source": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_without_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS\n)\n\nno_pack_val_loss = (job_without_sequence_packing_status.status_details or {}).get(\"val_loss\")\nif no_pack_val_loss is not None:\n print(f\"Validation loss: {float(no_pack_val_loss):.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")",
"language": "python",
- "source_html": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_without_sequence_packing.name,\n timeout=TIMEOUT_SECONDS\n)\n\nprint(f"Validation loss: {job_without_sequence_packing_status.status_details['val_loss']:.2f}")\n"
+ "source_html": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_without_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS\n)\n\nno_pack_val_loss = (job_without_sequence_packing_status.status_details or {}).get("val_loss")\nif no_pack_val_loss is not None:\n print(f"Validation loss: {float(no_pack_val_loss):.2f}")\nelse:\n print("Validation loss: not reported in job status")\n"
},
{
"type": "markdown",
@@ -143,9 +143,9 @@
},
{
"type": "code",
- "source": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = \"customization-training-job\"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n \"\"\"Calculate elapsed time in seconds from step's created_at to updated_at.\"\"\"\n created_at = datetime.fromisoformat(step.created_at.replace(\"Z\", \"+00:00\"))\n updated_at = datetime.fromisoformat(step.updated_at.replace(\"Z\", \"+00:00\"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_with_sequence_packing.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_without_sequence_packing.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n \"Seq Packing Enabled\": [True, False],\n \"Val Loss\": [\n job_with_sequence_packing_status.status_details['val_loss'],\n job_without_sequence_packing_status.status_details['val_loss']\n ],\n \"Training Step Time, sec\": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({\"Val Loss\": \"{:.2f}\", \"Training Step Time, sec\": \"{:.0f}\"}).hide(axis='index')",
+ "source": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = \"training\"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n \"\"\"Calculate elapsed time in seconds from step's created_at to updated_at.\"\"\"\n created_at = datetime.fromisoformat(step.created_at.replace(\"Z\", \"+00:00\"))\n updated_at = datetime.fromisoformat(step.updated_at.replace(\"Z\", \"+00:00\"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_with_sequence_packing.job.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_without_sequence_packing.job.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n \"Seq Packing Enabled\": [True, False],\n \"Val Loss\": [\n (job_with_sequence_packing_status.status_details or {}).get(\"val_loss\"),\n (job_without_sequence_packing_status.status_details or {}).get(\"val_loss\"),\n ],\n \"Training Step Time, sec\": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({\"Val Loss\": \"{:.2f}\", \"Training Step Time, sec\": \"{:.0f}\"}).hide(axis='index')",
"language": "python",
- "source_html": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = "customization-training-job"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n """Calculate elapsed time in seconds from step's created_at to updated_at."""\n created_at = datetime.fromisoformat(step.created_at.replace("Z", "+00:00"))\n updated_at = datetime.fromisoformat(step.updated_at.replace("Z", "+00:00"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_with_sequence_packing.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_without_sequence_packing.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n "Seq Packing Enabled": [True, False],\n "Val Loss": [\n job_with_sequence_packing_status.status_details['val_loss'],\n job_without_sequence_packing_status.status_details['val_loss']\n ],\n "Training Step Time, sec": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({"Val Loss": "{:.2f}", "Training Step Time, sec": "{:.0f}"}).hide(axis='index')\n"
+ "source_html": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = "training"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n """Calculate elapsed time in seconds from step's created_at to updated_at."""\n created_at = datetime.fromisoformat(step.created_at.replace("Z", "+00:00"))\n updated_at = datetime.fromisoformat(step.updated_at.replace("Z", "+00:00"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_with_sequence_packing.job.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_without_sequence_packing.job.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n "Seq Packing Enabled": [True, False],\n "Val Loss": [\n (job_with_sequence_packing_status.status_details or {}).get("val_loss"),\n (job_without_sequence_packing_status.status_details or {}).get("val_loss"),\n ],\n "Training Step Time, sec": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({"Val Loss": "{:.2f}", "Training Step Time, sec": "{:.0f}"}).hide(axis='index')\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/optimize-throughput.ts b/docs/fern/components/notebooks/optimize-throughput.ts
index f25f2d527f..59211fc597 100644
--- a/docs/fern/components/notebooks/optimize-throughput.ts
+++ b/docs/fern/components/notebooks/optimize-throughput.ts
@@ -1,7 +1,9 @@
-// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
-// SPDX-License-Identifier: Apache-2.0
-
-/** Auto-generated by ipynb-to-fern-json.py - do not edit */
+/**
+ * SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+ * SPDX-License-Identifier: Apache-2.0
+ *
+ * Auto-generated by ipynb-to-fern-json.py - do not edit manually.
+ */
export default { cells: [
{
"type": "markdown",
@@ -10,8 +12,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n
\n"
},
{
"type": "markdown",
@@ -53,9 +55,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
+ "source": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f\"{DATASET_PATH}/\", # Trailing slash uploads directory contents to fileset root\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
"language": "python",
- "source_html": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
+ "source_html": "# Create fileset to store SFT training data\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f"{DATASET_PATH}/", # Trailing slash uploads directory contents to fileset root\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
},
{
"type": "markdown",
@@ -81,14 +83,14 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "### 5. Create LoRA Job with Sequence Packing\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.",
- "source_html": "5. Create LoRA Job with Sequence Packing
\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.
\n"
+ "source": "### 5. Create LoRA Job with Sequence Packing\nCreate a LoRA customization job with **sequence packing** enabled via `AutomodelJobInput` (`batch.sequence_packing=True`).",
+ "source_html": "5. Create LoRA Job with Sequence Packing
\nCreate a LoRA customization job with sequence packing enabled via AutomodelJobInput (batch.sequence_packing=True).
\n"
},
{
"type": "code",
- "source": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n)\n\n# Enable sequence packing to improve throughput and GPU utilization\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\njob_spec = CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=SEQUENCE_PACKING_ENABLED,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_with_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=job_spec\n)\n\nprint(f\"Job ID: {job_with_sequence_packing.name}\")\nprint(f\"Output model: {job_with_sequence_packing.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"packing-job-{job_suffix}\"\nPACK_OUTPUT_NAME = f\"packing-out-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 4096,\n },\n schedule={\"epochs\": 1, \"val_check_interval\": 0.1},\n batch={\n \"global_batch_size\": 64,\n \"micro_batch_size\": 1,\n \"sequence_packing\": SEQUENCE_PACKING_ENABLED,\n },\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": PACK_OUTPUT_NAME},\n)\n\njob_with_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job_with_sequence_packing.job.name}\")\nprint(f\"Output adapter: {PACK_OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n)\n\n# Enable sequence packing to improve throughput and GPU utilization\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-sft-job-{job_suffix}"\njob_spec = CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=SEQUENCE_PACKING_ENABLED,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_with_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=job_spec\n)\n\nprint(f"Job ID: {job_with_sequence_packing.name}")\nprint(f"Output model: {job_with_sequence_packing.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\nSEQUENCE_PACKING_ENABLED = True\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"packing-job-{job_suffix}"\nPACK_OUTPUT_NAME = f"packing-out-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 4096,\n },\n schedule={"epochs": 1, "val_check_interval": 0.1},\n batch={\n "global_batch_size": 64,\n "micro_batch_size": 1,\n "sequence_packing": SEQUENCE_PACKING_ENABLED,\n },\n optimizer={"learning_rate": 5e-5},\n parallelism={"num_gpus_per_node": 1},\n output={"name": PACK_OUTPUT_NAME},\n)\n\njob_with_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job_with_sequence_packing.job.name}")\nprint(f"Output adapter: {PACK_OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -113,20 +115,20 @@ export default { cells: [
},
{
"type": "code",
- "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"loss\"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n val_loss = float(td[val_loss_key]) if val_loss_key in td else None\n train_loss = float(td[train_loss_key]) if train_loss_key in td else None\n break\n break\n\n # Fall back to top-level status_details\n if status.status_details:\n if val_loss is None and val_loss_key in status.status_details:\n val_loss = float(status.status_details[val_loss_key])\n if train_loss is None and train_loss_key in status.status_details:\n train_loss = float(status.status_details[train_loss_key])\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.name,\n timeout=TIMEOUT_SECONDS,\n)\n\nprint(f\"Validation loss: {job_with_sequence_packing_status.status_details['val_loss']:.2f}\")",
+ "source": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = \"val_loss\"\nTRAIN_LOSS_KEY = \"loss\"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n \"\"\"\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n \"\"\"\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f\"Timeout reached after {elapsed_min:.1f} minutes\"\n print(f\"\\n{error_message}\")\n print(\"Job did not complete within the timeout period.\")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get(\"phase\", \"\"))\n # Update current step if it's active or pending (not completed)\n if job_step.status in (\"active\", \"pending\"):\n current_step_name = job_step.name\n current_step_phase = phase or \"started\"\n\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td[\"step\"]) if \"step\" in td else None\n max_steps = cast(int, td[\"max_steps\"]) if \"max_steps\" in td else None\n training_phase = cast(str, td[\"phase\"]) if \"phase\" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f\"Step {step}/{max_steps} ({pct:.0f}%)\"\n if training_phase:\n step_str += f\" - {training_phase}\"\n else:\n if current_step_name and current_step_phase:\n step_str = f\"{current_step_name} - {current_step_phase}\"\n elif current_step_name:\n step_str = f\"{current_step_name}\"\n else:\n step_str = \"Waiting for training to start...\"\n elapsed_str = f\"Elapsed: {elapsed_min:.1f} min\"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == \"completed\":\n # Redraw dashboard one final time with \"completed\" status\n status_str = f\"Status: {status.status}\"\n if step is not None and max_steps is not None:\n step_str = f\"Step {max_steps}/{max_steps} (100%)\"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f\"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n return status\n elif status.status.lower() in (\"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n print(f\"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)\")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get(\"message\", \"\")\n if error_msg:\n print(f\"\\nError: {error_msg}\")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == \"error\":\n print(f\"\\nFailed step: {job_step.name}\")\n if job_step.error_details:\n step_error = job_step.error_details.get(\"message\", \"\")\n if step_error:\n print(f\"Step error: {step_error}\")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == \"error\" and hasattr(task, \"error_stack\") and task.error_stack:\n print(f\"\\nError stack trace:\\n{task.error_stack}\")\n elif task.status == \"error\" and task.error_details:\n task_error = task.error_details.get(\"message\", \"\")\n if task_error:\n print(f\"Task error: {task_error}\")\n break\n\n raise Exception(f\"Job finished with status: {status.status}\")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get(\"val_loss\")\nif packed_val_loss is not None:\n print(f\"Validation loss: {float(packed_val_loss):.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")",
"language": "python",
- "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "loss"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n val_loss = float(td[val_loss_key]) if val_loss_key in td else None\n train_loss = float(td[train_loss_key]) if train_loss_key in td else None\n break\n break\n\n # Fall back to top-level status_details\n if status.status_details:\n if val_loss is None and val_loss_key in status.status_details:\n val_loss = float(status.status_details[val_loss_key])\n if train_loss is None and train_loss_key in status.status_details:\n train_loss = float(status.status_details[train_loss_key])\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.name,\n timeout=TIMEOUT_SECONDS,\n)\n\nprint(f"Validation loss: {job_with_sequence_packing_status.status_details['val_loss']:.2f}")\n"
+ "source_html": "import time\nfrom typing import cast\nfrom IPython.display import clear_output\nfrom nemo_platform.types.shared import PlatformJobStatusResponse\n\n# Timeout set to 30 minutes to accommodate typical LoRA training duration for this dataset size.\n# Actual training time will vary based on hardware, model size, and dataset complexity.\nTIMEOUT_SECONDS = 30 * 60 # 30 minutes\nVAL_LOSS_KEY = "val_loss"\nTRAIN_LOSS_KEY = "loss"\n\n# ---------------------------------------------------------------------------\n# Job polling with live dashboard\n# ---------------------------------------------------------------------------\n\ndef wait_for_job(\n workspace: str,\n job_name: str,\n timeout: int = TIMEOUT_SECONDS,\n poll_interval: int = 10,\n val_loss_key: str = VAL_LOSS_KEY,\n train_loss_key: str = TRAIN_LOSS_KEY,\n) -> PlatformJobStatusResponse:\n """\n Poll job status until completed, failed, cancelled, or timeout.\n Displays a live dashboard with loss curves and GPU metrics.\n\n Args:\n workspace: The workspace where the job is running.\n job_name: The name of the job to monitor.\n timeout: Maximum time to wait in seconds (default: 30 minutes).\n poll_interval: Time between status checks in seconds (default: 10).\n\n Returns:\n The final job status response.\n """\n start_time = time.time()\n\n # Time-series accumulators required for plotting\n elapsed_mins: list[float] = []\n val_losses: list[float | None] = []\n train_losses: list[float | None] = []\n vram_history: list[list[float]] = []\n util_history: list[list[float]] = []\n\n while True:\n elapsed = time.time() - start_time\n elapsed_min = elapsed / 60\n\n # Check for timeout\n if elapsed > timeout:\n error_message = f"Timeout reached after {elapsed_min:.1f} minutes"\n print(f"\\n{error_message}")\n print("Job did not complete within the timeout period.")\n raise Exception(error_message)\n\n status = client.jobs.get_status(name=job_name, workspace=workspace)\n\n # -- Extract training progress from nested steps structure --\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n val_loss: float | None = None\n train_loss: float | None = None\n current_step_name: str | None = None\n current_step_phase: str | None = None\n\n for job_step in status.steps or []:\n # Track the current active step name and phase for progress display\n if job_step.tasks:\n task = job_step.tasks[0]\n td = task.status_details or {}\n phase = cast(str, td.get("phase", ""))\n # Update current step if it's active or pending (not completed)\n if job_step.status in ("active", "pending"):\n current_step_name = job_step.name\n current_step_phase = phase or "started"\n\n if job_step.name == "training":\n for task in job_step.tasks or []:\n td = task.status_details or {}\n step = cast(int, td["step"]) if "step" in td else None\n max_steps = cast(int, td["max_steps"]) if "max_steps" in td else None\n training_phase = cast(str, td["phase"]) if "phase" in td else None\n raw_val_loss = td.get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n raw_train_loss = td.get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n break\n break\n\n if val_loss is None:\n raw_val_loss = (status.status_details or {}).get(val_loss_key)\n val_loss = float(raw_val_loss) if raw_val_loss is not None else None\n if train_loss is None:\n raw_train_loss = (status.status_details or {}).get(train_loss_key)\n train_loss = float(raw_train_loss) if raw_train_loss is not None else None\n\n # -- Collect GPU snapshot --\n vram_pcts, util_pcts = _get_gpu_snapshot()\n\n # -- Append to accumulators used for the plots --\n elapsed_mins.append(elapsed_min)\n val_losses.append(val_loss)\n train_losses.append(train_loss)\n vram_history.append(vram_pcts)\n util_history.append(util_pcts)\n\n # -- Build status strings --\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n pct = step / max_steps * 100\n step_str = f"Step {step}/{max_steps} ({pct:.0f}%)"\n if training_phase:\n step_str += f" - {training_phase}"\n else:\n if current_step_name and current_step_phase:\n step_str = f"{current_step_name} - {current_step_phase}"\n elif current_step_name:\n step_str = f"{current_step_name}"\n else:\n step_str = "Waiting for training to start..."\n elapsed_str = f"Elapsed: {elapsed_min:.1f} min"\n\n # -- Redraw dashboard --\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n\n # -- Check terminal conditions --\n if status.status.lower() == "completed":\n # Redraw dashboard one final time with "completed" status\n status_str = f"Status: {status.status}"\n if step is not None and max_steps is not None:\n step_str = f"Step {max_steps}/{max_steps} (100%)"\n clear_output(wait=True)\n _draw_dashboard(\n elapsed_mins, val_losses, train_losses,\n vram_history, util_history,\n job_name, status_str, step_str, elapsed_str,\n )\n print(f"\\nJob completed in {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n return status\n elif status.status.lower() in ("failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n print(f"Total time elapsed: {elapsed_min:.1f} minutes ({elapsed:.0f}s)")\n\n # Print error details from the job level\n if status.error_details:\n error_msg = status.error_details.get("message", "")\n if error_msg:\n print(f"\\nError: {error_msg}")\n\n # Find and print error details from the failed step/task\n for job_step in status.steps or []:\n if job_step.status == "error":\n print(f"\\nFailed step: {job_step.name}")\n if job_step.error_details:\n step_error = job_step.error_details.get("message", "")\n if step_error:\n print(f"Step error: {step_error}")\n # Get error_stack from the failed task\n for task in job_step.tasks or []:\n if task.status == "error" and hasattr(task, "error_stack") and task.error_stack:\n print(f"\\nError stack trace:\\n{task.error_stack}")\n elif task.status == "error" and task.error_details:\n task_error = task.error_details.get("message", "")\n if task_error:\n print(f"Task error: {task_error}")\n break\n\n raise Exception(f"Job finished with status: {status.status}")\n\n time.sleep(poll_interval)\n\n\n# Wait for the job to complete\njob_with_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_with_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS,\n)\n\npacked_val_loss = (job_with_sequence_packing_status.status_details or {}).get("val_loss")\nif packed_val_loss is not None:\n print(f"Validation loss: {float(packed_val_loss):.2f}")\nelse:\n print("Validation loss: not reported in job status")\n"
},
{
"type": "markdown",
- "source": "### 7. Create LoRA Job without Sequence Packing\nCreate a customization job with sequence packing disabled. It's expected to take longer to complete.",
- "source_html": "7. Create LoRA Job without Sequence Packing
\nCreate a customization job with sequence packing disabled. It's expected to take longer to complete.
\n"
+ "source": "### 7. Create LoRA Job without Sequence Packing\nCreate a second Automodel LoRA job with `batch.sequence_packing=False` for comparison.",
+ "source_html": "7. Create LoRA Job without Sequence Packing
\nCreate a second Automodel LoRA job with batch.sequence_packing=False for comparison.
\n"
},
{
"type": "code",
- "source": "import uuid\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\n\njob_spec_no_packing = CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=False,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_without_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=job_spec_no_packing\n)\n\nprint(f\"Job ID: {job_without_sequence_packing.name}\")\nprint(f\"Output model: {job_without_sequence_packing.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"no-packing-job-{job_suffix}\"\nNO_PACK_OUTPUT_NAME = f\"no-packing-out-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 4096,\n },\n schedule={\"epochs\": 1, \"val_check_interval\": 0.1},\n batch={\n \"global_batch_size\": 64,\n \"micro_batch_size\": 1,\n \"sequence_packing\": False,\n },\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": NO_PACK_OUTPUT_NAME},\n)\n\njob_without_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job_without_sequence_packing.job.name}\")\nprint(f\"Output adapter: {NO_PACK_OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"my-sft-job-{job_suffix}"\n\njob_spec_no_packing = CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=1,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=4096,\n val_check_interval=0.1,\n micro_batch_size=1,\n sequence_packing=False,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n )\n)\n\njob_without_sequence_packing = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=job_spec_no_packing\n)\n\nprint(f"Job ID: {job_without_sequence_packing.name}")\nprint(f"Output model: {job_without_sequence_packing.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"no-packing-job-{job_suffix}"\nNO_PACK_OUTPUT_NAME = f"no-packing-out-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 4096,\n },\n schedule={"epochs": 1, "val_check_interval": 0.1},\n batch={\n "global_batch_size": 64,\n "micro_batch_size": 1,\n "sequence_packing": False,\n },\n optimizer={"learning_rate": 5e-5},\n parallelism={"num_gpus_per_node": 1},\n output={"name": NO_PACK_OUTPUT_NAME},\n)\n\njob_without_sequence_packing = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job_without_sequence_packing.job.name}")\nprint(f"Output adapter: {NO_PACK_OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -135,9 +137,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_without_sequence_packing.name,\n timeout=TIMEOUT_SECONDS\n)\n\nprint(f\"Validation loss: {job_without_sequence_packing_status.status_details['val_loss']:.2f}\")",
+ "source": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace=\"default\",\n job_name=job_without_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS\n)\n\nno_pack_val_loss = (job_without_sequence_packing_status.status_details or {}).get(\"val_loss\")\nif no_pack_val_loss is not None:\n print(f\"Validation loss: {float(no_pack_val_loss):.2f}\")\nelse:\n print(\"Validation loss: not reported in job status\")",
"language": "python",
- "source_html": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_without_sequence_packing.name,\n timeout=TIMEOUT_SECONDS\n)\n\nprint(f"Validation loss: {job_without_sequence_packing_status.status_details['val_loss']:.2f}")\n"
+ "source_html": "# Wait for the training step to complete\njob_without_sequence_packing_status = wait_for_job(\n workspace="default",\n job_name=job_without_sequence_packing.job.name,\n timeout=TIMEOUT_SECONDS\n)\n\nno_pack_val_loss = (job_without_sequence_packing_status.status_details or {}).get("val_loss")\nif no_pack_val_loss is not None:\n print(f"Validation loss: {float(no_pack_val_loss):.2f}")\nelse:\n print("Validation loss: not reported in job status")\n"
},
{
"type": "markdown",
@@ -146,9 +148,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = \"customization-training-job\"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n \"\"\"Calculate elapsed time in seconds from step's created_at to updated_at.\"\"\"\n created_at = datetime.fromisoformat(step.created_at.replace(\"Z\", \"+00:00\"))\n updated_at = datetime.fromisoformat(step.updated_at.replace(\"Z\", \"+00:00\"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_with_sequence_packing.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_without_sequence_packing.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n \"Seq Packing Enabled\": [True, False],\n \"Val Loss\": [\n job_with_sequence_packing_status.status_details['val_loss'],\n job_without_sequence_packing_status.status_details['val_loss']\n ],\n \"Training Step Time, sec\": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({\"Val Loss\": \"{:.2f}\", \"Training Step Time, sec\": \"{:.0f}\"}).hide(axis='index')",
+ "source": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = \"training\"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n \"\"\"Calculate elapsed time in seconds from step's created_at to updated_at.\"\"\"\n created_at = datetime.fromisoformat(step.created_at.replace(\"Z\", \"+00:00\"))\n updated_at = datetime.fromisoformat(step.updated_at.replace(\"Z\", \"+00:00\"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_with_sequence_packing.job.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace=\"default\",\n job=job_without_sequence_packing.job.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n \"Seq Packing Enabled\": [True, False],\n \"Val Loss\": [\n (job_with_sequence_packing_status.status_details or {}).get(\"val_loss\"),\n (job_without_sequence_packing_status.status_details or {}).get(\"val_loss\"),\n ],\n \"Training Step Time, sec\": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({\"Val Loss\": \"{:.2f}\", \"Training Step Time, sec\": \"{:.0f}\"}).hide(axis='index')",
"language": "python",
- "source_html": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = "customization-training-job"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n """Calculate elapsed time in seconds from step's created_at to updated_at."""\n created_at = datetime.fromisoformat(step.created_at.replace("Z", "+00:00"))\n updated_at = datetime.fromisoformat(step.updated_at.replace("Z", "+00:00"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_with_sequence_packing.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_without_sequence_packing.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n "Seq Packing Enabled": [True, False],\n "Val Loss": [\n job_with_sequence_packing_status.status_details['val_loss'],\n job_without_sequence_packing_status.status_details['val_loss']\n ],\n "Training Step Time, sec": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({"Val Loss": "{:.2f}", "Training Step Time, sec": "{:.0f}"}).hide(axis='index')\n"
+ "source_html": "from nemo_platform.types.jobs import PlatformJobStep\nfrom datetime import datetime\nimport pandas as pd\n\nSTEP_NAME = "training"\n\ndef get_elapsed_time(step: PlatformJobStep) -> float:\n """Calculate elapsed time in seconds from step's created_at to updated_at."""\n created_at = datetime.fromisoformat(step.created_at.replace("Z", "+00:00"))\n updated_at = datetime.fromisoformat(step.updated_at.replace("Z", "+00:00"))\n return (updated_at - created_at).total_seconds()\n\nstep_with_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_with_sequence_packing.job.name,\n)\n\nstep_without_sequence_packing = client.jobs.steps.retrieve(\n name=STEP_NAME,\n workspace="default",\n job=job_without_sequence_packing.job.name,\n)\n\ntime_to_complete_with_sequence_packing = get_elapsed_time(step_with_sequence_packing)\ntime_to_complete_without_sequence_packing = get_elapsed_time(step_without_sequence_packing)\n\n# Display results as a table\nresults_df = pd.DataFrame({\n "Seq Packing Enabled": [True, False],\n "Val Loss": [\n (job_with_sequence_packing_status.status_details or {}).get("val_loss"),\n (job_without_sequence_packing_status.status_details or {}).get("val_loss"),\n ],\n "Training Step Time, sec": [\n time_to_complete_with_sequence_packing,\n time_to_complete_without_sequence_packing\n ]\n})\n\nresults_df.style.format({"Val Loss": "{:.2f}", "Training Step Time, sec": "{:.0f}"}).hide(axis='index')\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/sft-customization-job.json b/docs/fern/components/notebooks/sft-customization-job.json
index 139e4d99d1..6356affdaa 100644
--- a/docs/fern/components/notebooks/sft-customization-job.json
+++ b/docs/fern/components/notebooks/sft-customization-job.json
@@ -7,8 +7,8 @@
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n
\n"
},
{
"type": "markdown",
@@ -79,9 +79,9 @@
},
{
"type": "code",
- "source": "# Create fileset to store SFT training data\nDATASET_NAME = \"sft-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
+ "source": "# Create fileset to store SFT training data\nDATASET_NAME = \"sft-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f\"{DATASET_PATH}/\", # Trailing slash uploads directory contents to fileset root\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
"language": "python",
- "source_html": "# Create fileset to store SFT training data\nDATASET_NAME = "sft-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
+ "source_html": "# Create fileset to store SFT training data\nDATASET_NAME = "sft-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f"{DATASET_PATH}/", # Trailing slash uploads directory contents to fileset root\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
},
{
"type": "markdown",
@@ -107,8 +107,8 @@
},
{
"type": "markdown",
- "source": "### 6. Create SFT Finetuning Job\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.",
- "source_html": "6. Create SFT Finetuning Job
\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.
\n"
+ "source": "### 6. Create SFT Finetuning Job\nCreate a customization job to fine-tune all model weights using the **Automodel** backend and `AutomodelJobInput`.",
+ "source_html": "6. Create SFT Finetuning Job
\nCreate a customization job to fine-tune all model weights using the Automodel backend and AutomodelJobInput.
\n"
},
{
"type": "markdown",
@@ -117,9 +117,9 @@
},
{
"type": "code",
- "source": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n )\n)\n\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\nOUTPUT_NAME = f\"sft-model-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"all_weights\",\n \"max_seq_length\": 2048,\n },\n schedule={\"epochs\": 2},\n batch={\"global_batch_size\": 64, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\n \"num_gpus_per_node\": 1,\n \"num_nodes\": 1,\n \"tensor_parallel_size\": 1,\n \"pipeline_parallel_size\": 1,\n },\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output model: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-sft-job-{job_suffix}"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n )\n)\n\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-sft-job-{job_suffix}"\nOUTPUT_NAME = f"sft-model-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "all_weights",\n "max_seq_length": 2048,\n },\n schedule={"epochs": 2},\n batch={"global_batch_size": 64, "micro_batch_size": 1},\n optimizer={"learning_rate": 5e-5},\n parallelism={\n "num_gpus_per_node": 1,\n "num_nodes": 1,\n "tensor_parallel_size": 1,\n "pipeline_parallel_size": 1,\n },\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output model: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -128,9 +128,9 @@
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace=\"default\"\n )\n\n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n\n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n\n time.sleep(10)",
+ "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n\n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n\n # Exit loop when job reaches a terminal status\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n\n time.sleep(10)\n\nif status.status != \"completed\":\n raise RuntimeError(f\"Training job finished with status: {status.status}\")",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace="default"\n )\n\n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n\n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n\n time.sleep(10)\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n\n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n\n # Exit loop when job reaches a terminal status\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n\n time.sleep(10)\n\nif status.status != "completed":\n raise RuntimeError(f"Training job finished with status: {status.status}")\n"
},
{
"type": "markdown",
@@ -144,25 +144,20 @@
},
{
"type": "code",
- "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))",
+ "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))",
"language": "python",
- "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))\n"
+ "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))\n"
},
{
"type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"sft-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"sft-model-deployment-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=\"nvcr.io/nim/nvidia/llm-nim\",\n image_tag=\"1.15.5\",\n gpu=1,\n model_name=job.spec.output.name, # ModelEntity name from training,\n model_namespace=\"default\", # Workspace where ModelEntity lives\n additional_envs={\"NIM_MODEL_PROFILE\": \"vllm\"}\n ),\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace=\"default\"\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {deployment_status.status}\")",
+ "source": "# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"sft-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"sft-model-deployment-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": OUTPUT_NAME,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n },\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace=\"default\"\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {deployment_status.status}\")",
"language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"sft-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"sft-model-deployment-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name="nvcr.io/nim/nvidia/llm-nim",\n image_tag="1.15.5",\n gpu=1,\n model_name=job.spec.output.name, # ModelEntity name from training,\n model_namespace="default", # Workspace where ModelEntity lives\n additional_envs={"NIM_MODEL_PROFILE": "vllm"}\n ),\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace="default"\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {deployment_status.status}")\n"
+ "source_html": "# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"sft-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"sft-model-deployment-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={\n "model_namespace": "default",\n "model_name": OUTPUT_NAME,\n },\n executor_config={\n "gpu": 1,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n },\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace="default"\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {deployment_status.status}")\n"
},
{
"type": "markdown",
- "source": "The deployment service automatically:\n- Downloads model weights from the Files service\n- Provisions storage (PVC) for the weights\n- Configures and starts the NIM container\n\n**Multi-GPU Deployment:**\n\nFor larger models requiring multiple GPUs, configure parallelism with environment variables:\n\n```python\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=\"sft-model-config-multigpu\",\n \n nim_deployment={\n \"image_name\": \"nvcr.io/nim/nvidia/llm-nim\",\n \"image_tag\": \"1.13.1\",\n \"gpu\": 2, # Total GPUs\n \"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"2\", # Tensor parallelism\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"1\" # Pipeline parallelism\n }\n }\n)\n```",
- "source_html": "The deployment service automatically:
\n\n- Downloads model weights from the Files service
\n- Provisions storage (PVC) for the weights
\n- Configures and starts the NIM container
\n
\nMulti-GPU Deployment:
\nFor larger models requiring multiple GPUs, configure parallelism with environment variables:
\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name="sft-model-config-multigpu",\n \n nim_deployment={\n "image_name": "nvcr.io/nim/nvidia/llm-nim",\n "image_tag": "1.13.1",\n "gpu": 2, # Total GPUs\n "additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "2", # Tensor parallelism\n "NIM_PIPELINE_PARALLEL_SIZE": "1" # Pipeline parallelism\n }\n }\n)\n
\n"
- },
- {
- "type": "markdown",
- "source": "**Single-Node Constraint:** Model deployments are limited to a single node. The maximum `gpu` value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.\n\n---\n\n#### GPU Parallelism\n\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the `NIM_TENSOR_PARALLEL_SIZE` and `NIM_PIPELINE_PARALLEL_SIZE` environment variables.\n\n| Strategy | Description | Best For |\n|----------|-------------|----------|\n| **Tensor Parallel (TP)** | Splits model layers across GPUs | Lowest latency |\n| **Pipeline Parallel (PP)** | Splits model depth across GPUs | Highest throughput |\n\n**Formula:** `gpu` = `NIM_TENSOR_PARALLEL_SIZE` × `NIM_PIPELINE_PARALLEL_SIZE`\n\n---\n\n#### Example Configurations\n\n**Default (TP=8, PP=1) — Lowest Latency**\n```\n\"gpu\": 8\n# NIM automatically sets NIM_TENSOR_PARALLEL_SIZE=8\n```\n\n**Balanced (TP=4, PP=2)**\n```\n\"gpu\": 8,\n\"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"4\",\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"2\"\n}\n```\n\n**Throughput Optimized (TP=2, PP=4)**\n```\n\"gpu\": 8,\n\"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"2\",\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"4\"\n}\n```",
- "source_html": "Single-Node Constraint: Model deployments are limited to a single node. The maximum gpu value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.
\n
\nGPU Parallelism
\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the NIM_TENSOR_PARALLEL_SIZE and NIM_PIPELINE_PARALLEL_SIZE environment variables.
\n\n\n\n| Strategy | \nDescription | \nBest For | \n
\n\n\n\n| Tensor Parallel (TP) | \nSplits model layers across GPUs | \nLowest latency | \n
\n\n| Pipeline Parallel (PP) | \nSplits model depth across GPUs | \nHighest throughput | \n
\n\n
\nFormula: gpu = NIM_TENSOR_PARALLEL_SIZE × NIM_PIPELINE_PARALLEL_SIZE
\n
\nExample Configurations
\nDefault (TP=8, PP=1) — Lowest Latency
\n"gpu": 8\n# NIM automatically sets NIM_TENSOR_PARALLEL_SIZE=8\n
\nBalanced (TP=4, PP=2)
\n"gpu": 8,\n"additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "4",\n "NIM_PIPELINE_PARALLEL_SIZE": "2"\n}\n
\nThroughput Optimized (TP=2, PP=4)
\n"gpu": 8,\n"additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "2",\n "NIM_PIPELINE_PARALLEL_SIZE": "4"\n}\n
\n"
+ "source": "The deployment service automatically:\n- Downloads model weights from the Files service\n- Provisions storage (PVC) for the weights\n- Configures and starts the vLLM container\n\n**Multi-GPU Deployment:**\n\nFor larger models requiring multiple GPUs, increase `gpu` in `executor_config`. vLLM computes tensor parallelism from the GPU count and model architecture:\n\n```python\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=\"sft-model-config-multigpu\",\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": OUTPUT_NAME,\n },\n executor_config={\n \"gpu\": 2,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n },\n)\n```\n\n**Single-Node Constraint:** Model deployments are limited to a single node. The maximum `gpu` value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.",
+ "source_html": "The deployment service automatically:
\n\n- Downloads model weights from the Files service
\n- Provisions storage (PVC) for the weights
\n- Configures and starts the vLLM container
\n
\nMulti-GPU Deployment:
\nFor larger models requiring multiple GPUs, increase gpu in executor_config. vLLM computes tensor parallelism from the GPU count and model architecture:
\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name="sft-model-config-multigpu",\n engine="vllm",\n model_spec={\n "model_namespace": "default",\n "model_name": OUTPUT_NAME,\n },\n executor_config={\n "gpu": 2,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n },\n)\n
\nSingle-Node Constraint: Model deployments are limited to a single node. The maximum gpu value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.
\n"
},
{
"type": "markdown",
@@ -182,14 +177,14 @@
},
{
"type": "code",
- "source": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\n\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\n\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": f\"default/{job.spec.output.name}\",\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 128\n }\n)\n\nprint(\"=\" * 60)\nprint(\"MODEL EVALUATION\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
+ "source": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\n\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\n\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": f\"default/{OUTPUT_NAME}\",\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 128\n }\n)\n\nprint(\"=\" * 60)\nprint(\"MODEL EVALUATION\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
"language": "python",
- "source_html": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\n\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\n\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={\n "model": f"default/{job.spec.output.name}",\n "messages": messages,\n "temperature": 0,\n "max_tokens": 128\n }\n)\n\nprint("=" * 60)\nprint("MODEL EVALUATION")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
+ "source_html": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\n\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\n\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={\n "model": f"default/{OUTPUT_NAME}",\n "messages": messages,\n "temperature": 0,\n "max_tokens": 128\n }\n)\n\nprint("=" * 60)\nprint("MODEL EVALUATION")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
},
{
"type": "markdown",
- "source": "#### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world examples from your use case\n- Compare responses to base model and expected outputs\n- Verify the model exhibits desired behavior changes\n- Check edge cases and error handling\n\n**What to look for:**\n- ✅ Model follows your desired output format\n- ✅ Applies domain knowledge correctly\n- ✅ Maintains general language capabilities\n- ✅ Avoids unwanted behaviors or biases\n- ❌ Doesn't hallucinate facts not in training data\n- ❌ Doesn't produce repetitive or nonsensical outputs\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n---\n\n\n## Troubleshooting\n\n**Job fails during model download:**\n- Verify authentication secrets are configured (refer to [Managing Secrets](../../get-started/concepts/manage-secrets.md))\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page\n- Check the `model_uri` format is correct (`fileset://`)\n- Ensure you have accepted the model's terms of service on HuggingFace\n- Check job status and logs: `client.customization.jobs.retrieve(name=job.name, workspace=\"default\")`\n\n**Job fails with OOM (Out of Memory) error:**\n1. **First try:** Reduce `micro_batch_size` from 2 to 1\n2. **Still OOM:** Reduce `batch_size` from 4 to 2\n3. **Still OOM:** Reduce `max_seq_length` from 2048 to 1024 or 512\n4. **Last resort:** Increase GPU count and use `tensor_parallel_size` for model sharding\n\n**Loss curves not decreasing (underfitting):**\n- Increase training duration: `epochs: 5-10` instead of 3\n- Adjust learning rate: Try `1e-5` to `1e-4`\n- Add warmup: Set `warmup_steps` to ~10% of total training steps\n- Check data quality: Verify formatting, remove duplicates, ensure diversity\n\n**Training loss decreases but validation loss increases (overfitting):**\n- Reduce epochs: Try `epochs: 1-2` instead of 5+\n- Lower learning rate: Use `2e-5` or `1e-5`\n- Increase dataset size and diversity\n- Verify train/validation split has no data leakage\n\n**Model output quality is poor despite good training metrics:**\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases\n- Review data quality, format, and diversity—metrics can be misleading with poor data\n- Try a different base model size or architecture\n- Adjust learning rate and batch size\n- Compare to baseline: Test base model to ensure fine-tuning improved performance\n\n**Deployment fails:**\n- Verify output model exists: `client.models.retrieve(name=job.spec.output.name, workspace=\"default\")`\n- Check deployment logs: `client.inference.deployments.get_logs(name=deployment.name, workspace=\"default\")`\n- Ensure sufficient GPU resources available for model size\n- Verify NIM image tag `1.13.1` is compatible with your model\n\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Learn about [LoRA customization](./lora-customization-job) for resource-efficient fine-tuning",
- "source_html": "Evaluation Best Practices
\nManual Evaluation (Recommended)
\n\n- Test with real-world examples from your use case
\n- Compare responses to base model and expected outputs
\n- Verify the model exhibits desired behavior changes
\n- Check edge cases and error handling
\n
\nWhat to look for:
\n\n- ✅ Model follows your desired output format
\n- ✅ Applies domain knowledge correctly
\n- ✅ Maintains general language capabilities
\n- ✅ Avoids unwanted behaviors or biases
\n- ❌ Doesn't hallucinate facts not in training data
\n- ❌ Doesn't produce repetitive or nonsensical outputs
\n
\n
\nHyperparameters
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\n
\nTroubleshooting
\nJob fails during model download:
\n\n- Verify authentication secrets are configured (refer to Managing Secrets)
\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page
\n- Check the
model_uri format is correct (fileset://) \n- Ensure you have accepted the model's terms of service on HuggingFace
\n- Check job status and logs:
client.customization.jobs.retrieve(name=job.name, workspace="default") \n
\nJob fails with OOM (Out of Memory) error:
\n\n- First try: Reduce
micro_batch_size from 2 to 1 \n- Still OOM: Reduce
batch_size from 4 to 2 \n- Still OOM: Reduce
max_seq_length from 2048 to 1024 or 512 \n- Last resort: Increase GPU count and use
tensor_parallel_size for model sharding \n
\nLoss curves not decreasing (underfitting):
\n\n- Increase training duration:
epochs: 5-10 instead of 3 \n- Adjust learning rate: Try
1e-5 to 1e-4 \n- Add warmup: Set
warmup_steps to ~10% of total training steps \n- Check data quality: Verify formatting, remove duplicates, ensure diversity
\n
\nTraining loss decreases but validation loss increases (overfitting):
\n\n- Reduce epochs: Try
epochs: 1-2 instead of 5+ \n- Lower learning rate: Use
2e-5 or 1e-5 \n- Increase dataset size and diversity
\n- Verify train/validation split has no data leakage
\n
\nModel output quality is poor despite good training metrics:
\n\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases
\n- Review data quality, format, and diversity—metrics can be misleading with poor data
\n- Try a different base model size or architecture
\n- Adjust learning rate and batch size
\n- Compare to baseline: Test base model to ensure fine-tuning improved performance
\n
\nDeployment fails:
\n\n- Verify output model exists:
client.models.retrieve(name=job.spec.output.name, workspace="default") \n- Check deployment logs:
client.inference.deployments.get_logs(name=deployment.name, workspace="default") \n- Ensure sufficient GPU resources available for model size
\n- Verify NIM image tag
1.13.1 is compatible with your model \n
\nNext Steps
\n\n"
+ "source": "#### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world examples from your use case\n- Compare responses to base model and expected outputs\n- Verify the model exhibits desired behavior changes\n- Check edge cases and error handling\n\n**What to look for:**\n- ✅ Model follows your desired output format\n- ✅ Applies domain knowledge correctly\n- ✅ Maintains general language capabilities\n- ✅ Avoids unwanted behaviors or biases\n- ❌ Doesn't hallucinate facts not in training data\n- ❌ Doesn't produce repetitive or nonsensical outputs\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n---\n\n\n## Troubleshooting\n\n**Job fails during model download:**\n- Verify authentication secrets are configured (refer to [Managing Secrets](../../get-started/concepts/manage-secrets.md))\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page (for example, [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct))\n- Confirm the model fileset uses `token_secret=hf_secret.name` for gated models\n- Check `AutomodelJobInput` references use the `workspace/name` format: `model=f\"default/{MODEL_NAME}\"` and `dataset={\"training\": f\"default/{DATASET_NAME}\"}` (for example, `default/llama-3-2-1b-base`, `default/sft-dataset`)\n- Verify the model entity points at the fileset: `fileset=f\"default/{MODEL_NAME}\"`\n- Check job status: `client.jobs.get_status(name=job.job.name, workspace=\"default\")`\n\n**Job fails with OOM (Out of Memory) error:**\n1. **First try:** Reduce `global_batch_size` from 64 to 32 or 16 in `batch={...}`\n2. **Still OOM:** Keep `micro_batch_size` at 1 (already the minimum in this tutorial)\n3. **Still OOM:** Reduce `max_seq_length` from 2048 to 1024 or 512 in `training={...}`\n4. **Last resort:** Increase `num_gpus_per_node` and `tensor_parallel_size` in `parallelism={...}`\n\n**Loss curves not decreasing (underfitting):**\n- Increase training duration: raise `epochs` from 2 to 3-5 in `schedule={...}`\n- Adjust learning rate: try `1e-4` or `1e-5` instead of the default `5e-5` in `optimizer={...}`\n- Check data quality: Verify formatting, remove duplicates, ensure diversity\n\n**Training loss decreases but validation loss increases (overfitting):**\n- Reduce `epochs` from 2 to 1 in `schedule={...}`\n- Lower `learning_rate` from `5e-5` to `2e-5` or `1e-5` in `optimizer={...}`\n- Increase dataset size and diversity\n- Verify train/validation split has no data leakage\n\n**Model output quality is poor despite good training metrics:**\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases\n- Review data quality, format, and diversity—metrics can be misleading with poor data\n- Try a different base model size or architecture\n- Adjust `learning_rate` and `global_batch_size`\n- Compare to baseline: Test base model to ensure fine-tuning improved performance\n\n**Deployment fails:**\n- Verify output model exists: `client.models.retrieve(name=OUTPUT_NAME, workspace=\"default\")`\n- Check deployment logs: `client.inference.deployments.get_logs(name=deployment.name, workspace=\"default\")`\n- Ensure sufficient GPU resources for `executor_config={\"gpu\": 1, ...}`\n- Verify the deployment config matches this tutorial: `engine=\"vllm\"` with `vllm/vllm-openai:v0.22.1`\n\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Learn about [LoRA customization](./lora-customization-job) for resource-efficient fine-tuning",
+ "source_html": "Evaluation Best Practices
\nManual Evaluation (Recommended)
\n\n- Test with real-world examples from your use case
\n- Compare responses to base model and expected outputs
\n- Verify the model exhibits desired behavior changes
\n- Check edge cases and error handling
\n
\nWhat to look for:
\n\n- ✅ Model follows your desired output format
\n- ✅ Applies domain knowledge correctly
\n- ✅ Maintains general language capabilities
\n- ✅ Avoids unwanted behaviors or biases
\n- ❌ Doesn't hallucinate facts not in training data
\n- ❌ Doesn't produce repetitive or nonsensical outputs
\n
\n
\nHyperparameters
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\n
\nTroubleshooting
\nJob fails during model download:
\n\n- Verify authentication secrets are configured (refer to Managing Secrets)
\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page (for example, meta-llama/Llama-3.2-1B-Instruct)
\n- Confirm the model fileset uses
token_secret=hf_secret.name for gated models \n- Check
AutomodelJobInput references use the workspace/name format: model=f"default/{MODEL_NAME}" and dataset={"training": f"default/{DATASET_NAME}"} (for example, default/llama-3-2-1b-base, default/sft-dataset) \n- Verify the model entity points at the fileset:
fileset=f"default/{MODEL_NAME}" \n- Check job status:
client.jobs.get_status(name=job.job.name, workspace="default") \n
\nJob fails with OOM (Out of Memory) error:
\n\n- First try: Reduce
global_batch_size from 64 to 32 or 16 in batch={...} \n- Still OOM: Keep
micro_batch_size at 1 (already the minimum in this tutorial) \n- Still OOM: Reduce
max_seq_length from 2048 to 1024 or 512 in training={...} \n- Last resort: Increase
num_gpus_per_node and tensor_parallel_size in parallelism={...} \n
\nLoss curves not decreasing (underfitting):
\n\n- Increase training duration: raise
epochs from 2 to 3-5 in schedule={...} \n- Adjust learning rate: try
1e-4 or 1e-5 instead of the default 5e-5 in optimizer={...} \n- Check data quality: Verify formatting, remove duplicates, ensure diversity
\n
\nTraining loss decreases but validation loss increases (overfitting):
\n\n- Reduce
epochs from 2 to 1 in schedule={...} \n- Lower
learning_rate from 5e-5 to 2e-5 or 1e-5 in optimizer={...} \n- Increase dataset size and diversity
\n- Verify train/validation split has no data leakage
\n
\nModel output quality is poor despite good training metrics:
\n\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases
\n- Review data quality, format, and diversity—metrics can be misleading with poor data
\n- Try a different base model size or architecture
\n- Adjust
learning_rate and global_batch_size \n- Compare to baseline: Test base model to ensure fine-tuning improved performance
\n
\nDeployment fails:
\n\n- Verify output model exists:
client.models.retrieve(name=OUTPUT_NAME, workspace="default") \n- Check deployment logs:
client.inference.deployments.get_logs(name=deployment.name, workspace="default") \n- Ensure sufficient GPU resources for
executor_config={"gpu": 1, ...} \n- Verify the deployment config matches this tutorial:
engine="vllm" with vllm/vllm-openai:v0.22.1 \n
\nNext Steps
\n\n"
}
]
}
\ No newline at end of file
diff --git a/docs/fern/components/notebooks/sft-customization-job.ts b/docs/fern/components/notebooks/sft-customization-job.ts
index 04876dce11..850b96b06a 100644
--- a/docs/fern/components/notebooks/sft-customization-job.ts
+++ b/docs/fern/components/notebooks/sft-customization-job.ts
@@ -1,7 +1,9 @@
-// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
-// SPDX-License-Identifier: Apache-2.0
-
-/** Auto-generated by ipynb-to-fern-json.py - do not edit */
+/**
+ * SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+ * SPDX-License-Identifier: Apache-2.0
+ *
+ * Auto-generated by ipynb-to-fern-json.py - do not edit manually.
+ */
export default { cells: [
{
"type": "markdown",
@@ -10,8 +12,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (included with `pip install nemo-platform`)",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (included with
pip install nemo-platform) \n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Completed the [Quickstart](../../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n2. **Installed the Python SDK** (PyPI wrapper: `pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- Installed the Python SDK (PyPI wrapper:
pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n
\n"
},
{
"type": "markdown",
@@ -82,9 +84,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Create fileset to store SFT training data\nDATASET_NAME = \"sft-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
+ "source": "# Create fileset to store SFT training data\nDATASET_NAME = \"sft-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=\"default\",\n name=DATASET_NAME,\n description=\"SFT training data\"\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f\"{DATASET_PATH}/\", # Trailing slash uploads directory contents to fileset root\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=\"default\"\n)\n\n# Validate training data is uploaded correctly\nprint(\"Training data:\")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=\"default\").data], indent=2))",
"language": "python",
- "source_html": "# Create fileset to store SFT training data\nDATASET_NAME = "sft-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=DATASET_PATH, # Local directory with your JSONL files\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
+ "source_html": "# Create fileset to store SFT training data\nDATASET_NAME = "sft-dataset"\n\ntry:\n client.files.filesets.create(\n workspace="default",\n name=DATASET_NAME,\n description="SFT training data"\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\n# Upload training data files individually to ensure correct structure\nclient.files.upload(\n local_path=f"{DATASET_PATH}/", # Trailing slash uploads directory contents to fileset root\n remote_path="",\n fileset=DATASET_NAME,\n workspace="default"\n)\n\n# Validate training data is uploaded correctly\nprint("Training data:")\nprint(json.dumps([f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace="default").data], indent=2))\n"
},
{
"type": "markdown",
@@ -110,8 +112,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "### 6. Create SFT Finetuning Job\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.",
- "source_html": "6. Create SFT Finetuning Job
\nCreate a customization job with an inline target referencing the base model and dataset filesets created in previous steps.
\n"
+ "source": "### 6. Create SFT Finetuning Job\nCreate a customization job to fine-tune all model weights using the **Automodel** backend and `AutomodelJobInput`.",
+ "source_html": "6. Create SFT Finetuning Job
\nCreate a customization job to fine-tune all model weights using the Automodel backend and AutomodelJobInput.
\n"
},
{
"type": "markdown",
@@ -120,9 +122,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace=\"default\",\n spec=CustomizationJobInputParam(\n model=f\"default/{base_model.name}\",\n dataset=f\"fileset://default/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n )\n)\n\nprint(f\"Job ID: {job.name}\")\nprint(f\"Output model: {job.spec.output.name}\")",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f\"my-sft-job-{job_suffix}\"\nOUTPUT_NAME = f\"sft-model-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"default/{base_model.name}\",\n dataset={\"training\": f\"default/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"all_weights\",\n \"max_seq_length\": 2048,\n },\n schedule={\"epochs\": 2},\n batch={\"global_batch_size\": 64, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": 5e-5},\n parallelism={\n \"num_gpus_per_node\": 1,\n \"num_nodes\": 1,\n \"tensor_parallel_size\": 1,\n \"pipeline_parallel_size\": 1,\n },\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=\"default\", name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output model: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-sft-job-{job_suffix}"\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n workspace="default",\n spec=CustomizationJobInputParam(\n model=f"default/{base_model.name}",\n dataset=f"fileset://default/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=2,\n batch_size=64,\n learning_rate=0.00005,\n max_seq_length=2048,\n micro_batch_size=1,\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n )\n)\n\nprint(f"Job ID: {job.name}")\nprint(f"Output model: {job.spec.output.name}")\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\n\nJOB_NAME = f"my-sft-job-{job_suffix}"\nOUTPUT_NAME = f"sft-model-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"default/{base_model.name}",\n dataset={"training": f"default/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "all_weights",\n "max_seq_length": 2048,\n },\n schedule={"epochs": 2},\n batch={"global_batch_size": 64, "micro_batch_size": 1},\n optimizer={"learning_rate": 5e-5},\n parallelism={\n "num_gpus_per_node": 1,\n "num_nodes": 1,\n "tensor_parallel_size": 1,\n "pipeline_parallel_size": 1,\n },\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace="default", name=JOB_NAME\n)\n\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output model: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -131,9 +133,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace=\"default\"\n )\n\n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"customization-training-job\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n\n # Exit loop when job is completed (or failed/cancelled)\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n\n time.sleep(10)",
+ "source": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace=\"default\"\n )\n\n clear_output(wait=True)\n print(f\"Job Status: {status.model_dump_json(indent=2)}\")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == \"training\":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get(\"step\")\n max_steps = task_details.get(\"max_steps\")\n training_phase = task_details.get(\"phase\")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f\"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)\")\n if training_phase:\n print(f\"Training Phase: {training_phase}\")\n else:\n print(\"Training step not started yet or progress info not available\")\n\n # Exit loop when job reaches a terminal status\n if status.status in (\"completed\", \"failed\", \"cancelled\", \"error\"):\n print(f\"\\nJob finished with status: {status.status}\")\n break\n\n time.sleep(10)\n\nif status.status != \"completed\":\n raise RuntimeError(f\"Training job finished with status: {status.status}\")",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.customization.jobs.get_status(\n name=job.name,\n workspace="default"\n )\n\n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "customization-training-job":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n\n # Exit loop when job is completed (or failed/cancelled)\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n\n time.sleep(10)\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\n# Poll job status every 10 seconds until completed\nwhile True:\n status = client.jobs.get_status(\n name=job.job.name,\n workspace="default"\n )\n\n clear_output(wait=True)\n print(f"Job Status: {status.model_dump_json(indent=2)}")\n\n # Extract training progress from nested steps structure\n step: int | None = None\n max_steps: int | None = None\n training_phase: str | None = None\n\n for job_step in status.steps or []:\n if job_step.name == "training":\n for task in job_step.tasks or []:\n task_details = task.status_details or {}\n step = task_details.get("step")\n max_steps = task_details.get("max_steps")\n training_phase = task_details.get("phase")\n break\n break\n\n if step is not None and max_steps is not None:\n progress_pct = (step / max_steps) * 100\n print(f"Training Progress: Step {step}/{max_steps} ({progress_pct:.1f}%)")\n if training_phase:\n print(f"Training Phase: {training_phase}")\n else:\n print("Training step not started yet or progress info not available")\n\n # Exit loop when job reaches a terminal status\n if status.status in ("completed", "failed", "cancelled", "error"):\n print(f"\\nJob finished with status: {status.status}")\n break\n\n time.sleep(10)\n\nif status.status != "completed":\n raise RuntimeError(f"Training job finished with status: {status.status}")\n"
},
{
"type": "markdown",
@@ -147,25 +149,20 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))",
+ "source": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))",
"language": "python",
- "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=job.spec.output.name)\nprint(model_entity.model_dump_json(indent=2))\n"
+ "source_html": "# Validate model entity exists\nmodel_entity = client.models.retrieve(workspace='default', name=OUTPUT_NAME)\nprint(model_entity.model_dump_json(indent=2))\n"
},
{
"type": "code",
- "source": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"sft-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"sft-model-deployment-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name=\"nvcr.io/nim/nvidia/llm-nim\",\n image_tag=\"1.15.5\",\n gpu=1,\n model_name=job.spec.output.name, # ModelEntity name from training,\n model_namespace=\"default\", # Workspace where ModelEntity lives\n additional_envs={\"NIM_MODEL_PROFILE\": \"vllm\"}\n ),\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace=\"default\"\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {deployment_status.status}\")",
+ "source": "# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"sft-model-deployment-cfg-{deploy_suffix}\"\nDEPLOYMENT_NAME = f\"sft-model-deployment-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": OUTPUT_NAME,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n },\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace=\"default\",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace=\"default\"\n)\n\nprint(f\"Deployment name: {deployment.name}\")\nprint(f\"Deployment status: {deployment_status.status}\")",
"language": "python",
- "source_html": "from nemo_platform.types.inference import NIMDeploymentParam\n\n# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"sft-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"sft-model-deployment-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n nim_deployment=NIMDeploymentParam(\n image_name="nvcr.io/nim/nvidia/llm-nim",\n image_tag="1.15.5",\n gpu=1,\n model_name=job.spec.output.name, # ModelEntity name from training,\n model_namespace="default", # Workspace where ModelEntity lives\n additional_envs={"NIM_MODEL_PROFILE": "vllm"}\n ),\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace="default"\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {deployment_status.status}")\n"
+ "source_html": "# Create deployment config\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"sft-model-deployment-cfg-{deploy_suffix}"\nDEPLOYMENT_NAME = f"sft-model-deployment-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={\n "model_namespace": "default",\n "model_name": OUTPUT_NAME,\n },\n executor_config={\n "gpu": 1,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n },\n)\n\n# Deploy model using deployment_config created above\ndeployment = client.inference.deployments.create(\n workspace="default",\n name=DEPLOYMENT_NAME,\n config=deployment_config.name\n)\n\n\n# Check deployment status\ndeployment_status = client.inference.deployments.retrieve(\n name=deployment.name,\n workspace="default"\n)\n\nprint(f"Deployment name: {deployment.name}")\nprint(f"Deployment status: {deployment_status.status}")\n"
},
{
"type": "markdown",
- "source": "The deployment service automatically:\n- Downloads model weights from the Files service\n- Provisions storage (PVC) for the weights\n- Configures and starts the NIM container\n\n**Multi-GPU Deployment:**\n\nFor larger models requiring multiple GPUs, configure parallelism with environment variables:\n\n```python\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=\"sft-model-config-multigpu\",\n \n nim_deployment={\n \"image_name\": \"nvcr.io/nim/nvidia/llm-nim\",\n \"image_tag\": \"1.13.1\",\n \"gpu\": 2, # Total GPUs\n \"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"2\", # Tensor parallelism\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"1\" # Pipeline parallelism\n }\n }\n)\n```",
- "source_html": "The deployment service automatically:
\n\n- Downloads model weights from the Files service
\n- Provisions storage (PVC) for the weights
\n- Configures and starts the NIM container
\n
\nMulti-GPU Deployment:
\nFor larger models requiring multiple GPUs, configure parallelism with environment variables:
\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name="sft-model-config-multigpu",\n \n nim_deployment={\n "image_name": "nvcr.io/nim/nvidia/llm-nim",\n "image_tag": "1.13.1",\n "gpu": 2, # Total GPUs\n "additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "2", # Tensor parallelism\n "NIM_PIPELINE_PARALLEL_SIZE": "1" # Pipeline parallelism\n }\n }\n)\n
\n"
- },
- {
- "type": "markdown",
- "source": "**Single-Node Constraint:** Model deployments are limited to a single node. The maximum `gpu` value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.\n\n---\n\n#### GPU Parallelism\n\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the `NIM_TENSOR_PARALLEL_SIZE` and `NIM_PIPELINE_PARALLEL_SIZE` environment variables.\n\n| Strategy | Description | Best For |\n|----------|-------------|----------|\n| **Tensor Parallel (TP)** | Splits model layers across GPUs | Lowest latency |\n| **Pipeline Parallel (PP)** | Splits model depth across GPUs | Highest throughput |\n\n**Formula:** `gpu` = `NIM_TENSOR_PARALLEL_SIZE` × `NIM_PIPELINE_PARALLEL_SIZE`\n\n---\n\n#### Example Configurations\n\n**Default (TP=8, PP=1) — Lowest Latency**\n```\n\"gpu\": 8\n# NIM automatically sets NIM_TENSOR_PARALLEL_SIZE=8\n```\n\n**Balanced (TP=4, PP=2)**\n```\n\"gpu\": 8,\n\"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"4\",\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"2\"\n}\n```\n\n**Throughput Optimized (TP=2, PP=4)**\n```\n\"gpu\": 8,\n\"additional_envs\": {\n \"NIM_TENSOR_PARALLEL_SIZE\": \"2\",\n \"NIM_PIPELINE_PARALLEL_SIZE\": \"4\"\n}\n```",
- "source_html": "Single-Node Constraint: Model deployments are limited to a single node. The maximum gpu value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.
\n
\nGPU Parallelism
\nBy default, NIM uses all GPUs for tensor parallelism (TP). You can customize this behavior using the NIM_TENSOR_PARALLEL_SIZE and NIM_PIPELINE_PARALLEL_SIZE environment variables.
\n\n\n\n| Strategy | \nDescription | \nBest For | \n
\n\n\n\n| Tensor Parallel (TP) | \nSplits model layers across GPUs | \nLowest latency | \n
\n\n| Pipeline Parallel (PP) | \nSplits model depth across GPUs | \nHighest throughput | \n
\n\n
\nFormula: gpu = NIM_TENSOR_PARALLEL_SIZE × NIM_PIPELINE_PARALLEL_SIZE
\n
\nExample Configurations
\nDefault (TP=8, PP=1) — Lowest Latency
\n"gpu": 8\n# NIM automatically sets NIM_TENSOR_PARALLEL_SIZE=8\n
\nBalanced (TP=4, PP=2)
\n"gpu": 8,\n"additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "4",\n "NIM_PIPELINE_PARALLEL_SIZE": "2"\n}\n
\nThroughput Optimized (TP=2, PP=4)
\n"gpu": 8,\n"additional_envs": {\n "NIM_TENSOR_PARALLEL_SIZE": "2",\n "NIM_PIPELINE_PARALLEL_SIZE": "4"\n}\n
\n"
+ "source": "The deployment service automatically:\n- Downloads model weights from the Files service\n- Provisions storage (PVC) for the weights\n- Configures and starts the vLLM container\n\n**Multi-GPU Deployment:**\n\nFor larger models requiring multiple GPUs, increase `gpu` in `executor_config`. vLLM computes tensor parallelism from the GPU count and model architecture:\n\n```python\ndeployment_config = client.inference.deployment_configs.create(\n workspace=\"default\",\n name=\"sft-model-config-multigpu\",\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": \"default\",\n \"model_name\": OUTPUT_NAME,\n },\n executor_config={\n \"gpu\": 2,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n },\n)\n```\n\n**Single-Node Constraint:** Model deployments are limited to a single node. The maximum `gpu` value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.",
+ "source_html": "The deployment service automatically:
\n\n- Downloads model weights from the Files service
\n- Provisions storage (PVC) for the weights
\n- Configures and starts the vLLM container
\n
\nMulti-GPU Deployment:
\nFor larger models requiring multiple GPUs, increase gpu in executor_config. vLLM computes tensor parallelism from the GPU count and model architecture:
\ndeployment_config = client.inference.deployment_configs.create(\n workspace="default",\n name="sft-model-config-multigpu",\n engine="vllm",\n model_spec={\n "model_namespace": "default",\n "model_name": OUTPUT_NAME,\n },\n executor_config={\n "gpu": 2,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n },\n)\n
\nSingle-Node Constraint: Model deployments are limited to a single node. The maximum gpu value depends on the total GPUs available on a single node in your cluster. Multi-node deployments are not supported.
\n"
},
{
"type": "markdown",
@@ -185,13 +182,13 @@ export default { cells: [
},
{
"type": "code",
- "source": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\n\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\n\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": f\"default/{job.spec.output.name}\",\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 128\n }\n)\n\nprint(\"=\" * 60)\nprint(\"MODEL EVALUATION\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
+ "source": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = \"The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit.\"\nquestion = \"Who was the first person to walk on the Moon?\"\n\nmessages = [\n {\"role\": \"user\", \"content\": f\"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}\"}\n]\n\nresponse = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment.name,\n workspace=\"default\",\n body={\n \"model\": f\"default/{OUTPUT_NAME}\",\n \"messages\": messages,\n \"temperature\": 0,\n \"max_tokens\": 128\n }\n)\n\nprint(\"=\" * 60)\nprint(\"MODEL EVALUATION\")\nprint(\"=\" * 60)\nprint(f\"Question: {question}\")\nprint(f\"Expected: Neil Armstrong\")\nprint(f\"Model output: {response['choices'][0]['message']['content']}\")",
"language": "python",
- "source_html": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\n\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\n\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={\n "model": f"default/{job.spec.output.name}",\n "messages": messages,\n "temperature": 0,\n "max_tokens": 128\n }\n)\n\nprint("=" * 60)\nprint("MODEL EVALUATION")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
+ "source_html": "# Wait for deployment to be ready, then test\n# Test the fine-tuned model with a question answering prompt\ncontext = "The Apollo 11 mission was the first manned mission to land on the Moon. It was launched on July 16, 1969, and Neil Armstrong became the first person to walk on the lunar surface on July 20, 1969. Buzz Aldrin joined him shortly after, while Michael Collins remained in lunar orbit."\nquestion = "Who was the first person to walk on the Moon?"\n\nmessages = [\n {"role": "user", "content": f"Based on the following context, answer the question.\\n\\nContext: {context}\\n\\nQuestion: {question}"}\n]\n\nresponse = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment.name,\n workspace="default",\n body={\n "model": f"default/{OUTPUT_NAME}",\n "messages": messages,\n "temperature": 0,\n "max_tokens": 128\n }\n)\n\nprint("=" * 60)\nprint("MODEL EVALUATION")\nprint("=" * 60)\nprint(f"Question: {question}")\nprint(f"Expected: Neil Armstrong")\nprint(f"Model output: {response['choices'][0]['message']['content']}")\n"
},
{
"type": "markdown",
- "source": "#### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world examples from your use case\n- Compare responses to base model and expected outputs\n- Verify the model exhibits desired behavior changes\n- Check edge cases and error handling\n\n**What to look for:**\n- ✅ Model follows your desired output format\n- ✅ Applies domain knowledge correctly\n- ✅ Maintains general language capabilities\n- ✅ Avoids unwanted behaviors or biases\n- ❌ Doesn't hallucinate facts not in training data\n- ❌ Doesn't produce repetitive or nonsensical outputs\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n---\n\n\n## Troubleshooting\n\n**Job fails during model download:**\n- Verify authentication secrets are configured (refer to [Managing Secrets](../../get-started/concepts/manage-secrets.md))\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page\n- Check the `model_uri` format is correct (`fileset://`)\n- Ensure you have accepted the model's terms of service on HuggingFace\n- Check job status and logs: `client.customization.jobs.retrieve(name=job.name, workspace=\"default\")`\n\n**Job fails with OOM (Out of Memory) error:**\n1. **First try:** Reduce `micro_batch_size` from 2 to 1\n2. **Still OOM:** Reduce `batch_size` from 4 to 2\n3. **Still OOM:** Reduce `max_seq_length` from 2048 to 1024 or 512\n4. **Last resort:** Increase GPU count and use `tensor_parallel_size` for model sharding\n\n**Loss curves not decreasing (underfitting):**\n- Increase training duration: `epochs: 5-10` instead of 3\n- Adjust learning rate: Try `1e-5` to `1e-4`\n- Add warmup: Set `warmup_steps` to ~10% of total training steps\n- Check data quality: Verify formatting, remove duplicates, ensure diversity\n\n**Training loss decreases but validation loss increases (overfitting):**\n- Reduce epochs: Try `epochs: 1-2` instead of 5+\n- Lower learning rate: Use `2e-5` or `1e-5`\n- Increase dataset size and diversity\n- Verify train/validation split has no data leakage\n\n**Model output quality is poor despite good training metrics:**\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases\n- Review data quality, format, and diversity—metrics can be misleading with poor data\n- Try a different base model size or architecture\n- Adjust learning rate and batch size\n- Compare to baseline: Test base model to ensure fine-tuning improved performance\n\n**Deployment fails:**\n- Verify output model exists: `client.models.retrieve(name=job.spec.output.name, workspace=\"default\")`\n- Check deployment logs: `client.inference.deployments.get_logs(name=deployment.name, workspace=\"default\")`\n- Ensure sufficient GPU resources available for model size\n- Verify NIM image tag `1.13.1` is compatible with your model\n\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Learn about [LoRA customization](./lora-customization-job) for resource-efficient fine-tuning",
- "source_html": "Evaluation Best Practices
\nManual Evaluation (Recommended)
\n\n- Test with real-world examples from your use case
\n- Compare responses to base model and expected outputs
\n- Verify the model exhibits desired behavior changes
\n- Check edge cases and error handling
\n
\nWhat to look for:
\n\n- ✅ Model follows your desired output format
\n- ✅ Applies domain knowledge correctly
\n- ✅ Maintains general language capabilities
\n- ✅ Avoids unwanted behaviors or biases
\n- ❌ Doesn't hallucinate facts not in training data
\n- ❌ Doesn't produce repetitive or nonsensical outputs
\n
\n
\nHyperparameters
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\n
\nTroubleshooting
\nJob fails during model download:
\n\n- Verify authentication secrets are configured (refer to Managing Secrets)
\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page
\n- Check the
model_uri format is correct (fileset://) \n- Ensure you have accepted the model's terms of service on HuggingFace
\n- Check job status and logs:
client.customization.jobs.retrieve(name=job.name, workspace="default") \n
\nJob fails with OOM (Out of Memory) error:
\n\n- First try: Reduce
micro_batch_size from 2 to 1 \n- Still OOM: Reduce
batch_size from 4 to 2 \n- Still OOM: Reduce
max_seq_length from 2048 to 1024 or 512 \n- Last resort: Increase GPU count and use
tensor_parallel_size for model sharding \n
\nLoss curves not decreasing (underfitting):
\n\n- Increase training duration:
epochs: 5-10 instead of 3 \n- Adjust learning rate: Try
1e-5 to 1e-4 \n- Add warmup: Set
warmup_steps to ~10% of total training steps \n- Check data quality: Verify formatting, remove duplicates, ensure diversity
\n
\nTraining loss decreases but validation loss increases (overfitting):
\n\n- Reduce epochs: Try
epochs: 1-2 instead of 5+ \n- Lower learning rate: Use
2e-5 or 1e-5 \n- Increase dataset size and diversity
\n- Verify train/validation split has no data leakage
\n
\nModel output quality is poor despite good training metrics:
\n\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases
\n- Review data quality, format, and diversity—metrics can be misleading with poor data
\n- Try a different base model size or architecture
\n- Adjust learning rate and batch size
\n- Compare to baseline: Test base model to ensure fine-tuning improved performance
\n
\nDeployment fails:
\n\n- Verify output model exists:
client.models.retrieve(name=job.spec.output.name, workspace="default") \n- Check deployment logs:
client.inference.deployments.get_logs(name=deployment.name, workspace="default") \n- Ensure sufficient GPU resources available for model size
\n- Verify NIM image tag
1.13.1 is compatible with your model \n
\nNext Steps
\n\n"
+ "source": "#### Evaluation Best Practices\n\n**Manual Evaluation** (Recommended)\n- Test with real-world examples from your use case\n- Compare responses to base model and expected outputs\n- Verify the model exhibits desired behavior changes\n- Check edge cases and error handling\n\n**What to look for:**\n- ✅ Model follows your desired output format\n- ✅ Applies domain knowledge correctly\n- ✅ Maintains general language capabilities\n- ✅ Avoids unwanted behaviors or biases\n- ❌ Doesn't hallucinate facts not in training data\n- ❌ Doesn't produce repetitive or nonsensical outputs\n\n---\n\n## Hyperparameters\n\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the [Hyperparameter Reference](../manage-customization-jobs/hyperparameters.md).\n\n---\n\n\n## Troubleshooting\n\n**Job fails during model download:**\n- Verify authentication secrets are configured (refer to [Managing Secrets](../../get-started/concepts/manage-secrets.md))\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page (for example, [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct))\n- Confirm the model fileset uses `token_secret=hf_secret.name` for gated models\n- Check `AutomodelJobInput` references use the `workspace/name` format: `model=f\"default/{MODEL_NAME}\"` and `dataset={\"training\": f\"default/{DATASET_NAME}\"}` (for example, `default/llama-3-2-1b-base`, `default/sft-dataset`)\n- Verify the model entity points at the fileset: `fileset=f\"default/{MODEL_NAME}\"`\n- Check job status: `client.jobs.get_status(name=job.job.name, workspace=\"default\")`\n\n**Job fails with OOM (Out of Memory) error:**\n1. **First try:** Reduce `global_batch_size` from 64 to 32 or 16 in `batch={...}`\n2. **Still OOM:** Keep `micro_batch_size` at 1 (already the minimum in this tutorial)\n3. **Still OOM:** Reduce `max_seq_length` from 2048 to 1024 or 512 in `training={...}`\n4. **Last resort:** Increase `num_gpus_per_node` and `tensor_parallel_size` in `parallelism={...}`\n\n**Loss curves not decreasing (underfitting):**\n- Increase training duration: raise `epochs` from 2 to 3-5 in `schedule={...}`\n- Adjust learning rate: try `1e-4` or `1e-5` instead of the default `5e-5` in `optimizer={...}`\n- Check data quality: Verify formatting, remove duplicates, ensure diversity\n\n**Training loss decreases but validation loss increases (overfitting):**\n- Reduce `epochs` from 2 to 1 in `schedule={...}`\n- Lower `learning_rate` from `5e-5` to `2e-5` or `1e-5` in `optimizer={...}`\n- Increase dataset size and diversity\n- Verify train/validation split has no data leakage\n\n**Model output quality is poor despite good training metrics:**\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases\n- Review data quality, format, and diversity—metrics can be misleading with poor data\n- Try a different base model size or architecture\n- Adjust `learning_rate` and `global_batch_size`\n- Compare to baseline: Test base model to ensure fine-tuning improved performance\n\n**Deployment fails:**\n- Verify output model exists: `client.models.retrieve(name=OUTPUT_NAME, workspace=\"default\")`\n- Check deployment logs: `client.inference.deployments.get_logs(name=deployment.name, workspace=\"default\")`\n- Ensure sufficient GPU resources for `executor_config={\"gpu\": 1, ...}`\n- Verify the deployment config matches this tutorial: `engine=\"vllm\"` with `vllm/vllm-openai:v0.22.1`\n\n\n## Next Steps\n\n- [Monitor training metrics](fine-tune-metrics) in detail\n- [Evaluate your fine-tuned model](../../evaluator/index) using the Evaluator service\n- Learn about [LoRA customization](./lora-customization-job) for resource-efficient fine-tuning",
+ "source_html": "Evaluation Best Practices
\nManual Evaluation (Recommended)
\n\n- Test with real-world examples from your use case
\n- Compare responses to base model and expected outputs
\n- Verify the model exhibits desired behavior changes
\n- Check edge cases and error handling
\n
\nWhat to look for:
\n\n- ✅ Model follows your desired output format
\n- ✅ Applies domain knowledge correctly
\n- ✅ Maintains general language capabilities
\n- ✅ Avoids unwanted behaviors or biases
\n- ❌ Doesn't hallucinate facts not in training data
\n- ❌ Doesn't produce repetitive or nonsensical outputs
\n
\n
\nHyperparameters
\nFor detailed information on all available hyperparameters, recommended values, and tuning guidance, refer to the Hyperparameter Reference.
\n
\nTroubleshooting
\nJob fails during model download:
\n\n- Verify authentication secrets are configured (refer to Managing Secrets)
\n- For gated HuggingFace models (Llama, Gemma), accept the license on the model page (for example, meta-llama/Llama-3.2-1B-Instruct)
\n- Confirm the model fileset uses
token_secret=hf_secret.name for gated models \n- Check
AutomodelJobInput references use the workspace/name format: model=f"default/{MODEL_NAME}" and dataset={"training": f"default/{DATASET_NAME}"} (for example, default/llama-3-2-1b-base, default/sft-dataset) \n- Verify the model entity points at the fileset:
fileset=f"default/{MODEL_NAME}" \n- Check job status:
client.jobs.get_status(name=job.job.name, workspace="default") \n
\nJob fails with OOM (Out of Memory) error:
\n\n- First try: Reduce
global_batch_size from 64 to 32 or 16 in batch={...} \n- Still OOM: Keep
micro_batch_size at 1 (already the minimum in this tutorial) \n- Still OOM: Reduce
max_seq_length from 2048 to 1024 or 512 in training={...} \n- Last resort: Increase
num_gpus_per_node and tensor_parallel_size in parallelism={...} \n
\nLoss curves not decreasing (underfitting):
\n\n- Increase training duration: raise
epochs from 2 to 3-5 in schedule={...} \n- Adjust learning rate: try
1e-4 or 1e-5 instead of the default 5e-5 in optimizer={...} \n- Check data quality: Verify formatting, remove duplicates, ensure diversity
\n
\nTraining loss decreases but validation loss increases (overfitting):
\n\n- Reduce
epochs from 2 to 1 in schedule={...} \n- Lower
learning_rate from 5e-5 to 2e-5 or 1e-5 in optimizer={...} \n- Increase dataset size and diversity
\n- Verify train/validation split has no data leakage
\n
\nModel output quality is poor despite good training metrics:
\n\n- Training metrics optimize for loss, not your actual task—evaluate on real use cases
\n- Review data quality, format, and diversity—metrics can be misleading with poor data
\n- Try a different base model size or architecture
\n- Adjust
learning_rate and global_batch_size \n- Compare to baseline: Test base model to ensure fine-tuning improved performance
\n
\nDeployment fails:
\n\n- Verify output model exists:
client.models.retrieve(name=OUTPUT_NAME, workspace="default") \n- Check deployment logs:
client.inference.deployments.get_logs(name=deployment.name, workspace="default") \n- Ensure sufficient GPU resources for
executor_config={"gpu": 1, ...} \n- Verify the deployment config matches this tutorial:
engine="vllm" with vllm/vllm-openai:v0.22.1 \n
\nNext Steps
\n\n"
}
] };
diff --git a/docs/fern/components/notebooks/tool-calling.json b/docs/fern/components/notebooks/tool-calling.json
index d8af8b353c..0d7d76d052 100644
--- a/docs/fern/components/notebooks/tool-calling.json
+++ b/docs/fern/components/notebooks/tool-calling.json
@@ -7,8 +7,8 @@
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Installed the Python SDK with the Data Designer extra** (`uv pip install nemo-platform[data-designer]`)\n1. **Completed the [Quickstart](../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n1. (Optional if running outside of Quickstart) **Authenticated with the platform** using the CLI:\n ```bash\n nemo auth login\n ```\n For non-default URLs: `nemo auth login --base-url `\n1. Set your **Hugging Face token** as an environment variable before running the notebook (`export HF_TOKEN=\"\"`)\n1. **Accepted the dataset license** for [Salesforce/xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k) on Hugging Face (used for evaluation) ",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Installed the Python SDK with the Data Designer extra (
uv pip install nemo-platform[data-designer]) \n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- (Optional if running outside of Quickstart) Authenticated with the platform using the CLI:
nemo auth login\n
\nFor non-default URLs: nemo auth login --base-url <YOUR_NMP_BASE_URL> \n- Set your Hugging Face token as an environment variable before running the notebook (
export HF_TOKEN="<your-token>") \n- Accepted the dataset license for Salesforce/xlam-function-calling-60k on Hugging Face (used for evaluation)
\n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Installed the Python SDK with Data Designer support** (PyPI wrapper: `uv pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n1. **Completed the [Quickstart](../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n1. (Optional if running outside of Quickstart) **Authenticated with the platform** using the CLI:\n ```bash\n nemo auth login\n ```\n For non-default URLs: `nemo auth login --base-url `\n1. Set your **Hugging Face token** as an environment variable before running the notebook (`export HF_TOKEN=\"\"`)\n1. **Accepted the dataset license** for [Salesforce/xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k) on Hugging Face (used for evaluation) ",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Installed the Python SDK with Data Designer support (PyPI wrapper:
uv pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- (Optional if running outside of Quickstart) Authenticated with the platform using the CLI:
nemo auth login\n
\nFor non-default URLs: nemo auth login --base-url <YOUR_NMP_BASE_URL> \n- Set your Hugging Face token as an environment variable before running the notebook (
export HF_TOKEN="<your-token>") \n- Accepted the dataset license for Salesforce/xlam-function-calling-60k on Hugging Face (used for evaluation)
\n
\n"
},
{
"type": "markdown",
@@ -138,9 +138,9 @@
},
{
"type": "code",
- "source": "DATASET_NAME = \"tool-calling-dataset\"\n\ntry:\n client.files.filesets.create(\n name=DATASET_NAME,\n description=\"synthetic tool-calling training and validation data in OpenAI chat format\",\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path=\"\",\n fileset=DATASET_NAME,\n)\n\nprint(\"\\nTraining data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME).data],\n indent=2,\n))",
+ "source": "DATASET_NAME = \"tool-calling-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=DATASET_NAME,\n description=\"synthetic tool-calling training and validation data in OpenAI chat format\",\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint(\"\\nTraining data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))",
"language": "python",
- "source_html": "DATASET_NAME = "tool-calling-dataset"\n\ntry:\n client.files.filesets.create(\n name=DATASET_NAME,\n description="synthetic tool-calling training and validation data in OpenAI chat format",\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path="",\n fileset=DATASET_NAME,\n)\n\nprint("\\nTraining data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME).data],\n indent=2,\n))\n"
+ "source_html": "DATASET_NAME = "tool-calling-dataset"\n\ntry:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=DATASET_NAME,\n description="synthetic tool-calling training and validation data in OpenAI chat format",\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path="",\n fileset=DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint("\\nTraining data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n"
},
{
"type": "markdown",
@@ -149,9 +149,9 @@
},
{
"type": "code",
- "source": "try:\n client.files.filesets.create(\n name=EVAL_DATASET_NAME,\n description=\"synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)\",\n )\n print(f\"Created fileset: {EVAL_DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path=\"\",\n fileset=EVAL_DATASET_NAME,\n)\n\nprint(\"\\nEvaluation data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME).data],\n indent=2,\n))",
+ "source": "try:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=EVAL_DATASET_NAME,\n description=\"synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)\",\n )\n print(f\"Created fileset: {EVAL_DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path=\"\",\n fileset=EVAL_DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint(\"\\nEvaluation data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))",
"language": "python",
- "source_html": "try:\n client.files.filesets.create(\n name=EVAL_DATASET_NAME,\n description="synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)",\n )\n print(f"Created fileset: {EVAL_DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path="",\n fileset=EVAL_DATASET_NAME,\n)\n\nprint("\\nEvaluation data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME).data],\n indent=2,\n))\n"
+ "source_html": "try:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=EVAL_DATASET_NAME,\n description="synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)",\n )\n print(f"Created fileset: {EVAL_DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path="",\n fileset=EVAL_DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint("\\nEvaluation data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n"
},
{
"type": "markdown",
@@ -160,9 +160,9 @@
},
{
"type": "code",
- "source": "HF_TOKEN = os.environ.get(\"HF_TOKEN\")\nif HF_TOKEN is None:\n raise ValueError(\"HF_TOKEN is not set.\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name)\n\n\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(hf_secret.model_dump_json(indent=2))",
+ "source": "HF_TOKEN = os.environ.get(\"HF_TOKEN\")\nif HF_TOKEN is None:\n raise ValueError(\"HF_TOKEN is not set.\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n workspace=WORKSPACE,\n name=name,\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=WORKSPACE)\n\n\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(hf_secret.model_dump_json(indent=2))",
"language": "python",
- "source_html": "HF_TOKEN = os.environ.get("HF_TOKEN")\nif HF_TOKEN is None:\n raise ValueError("HF_TOKEN is not set.")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name)\n\n\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(hf_secret.model_dump_json(indent=2))\n"
+ "source_html": "HF_TOKEN = os.environ.get("HF_TOKEN")\nif HF_TOKEN is None:\n raise ValueError("HF_TOKEN is not set.")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n workspace=WORKSPACE,\n name=name,\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace=WORKSPACE)\n\n\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(hf_secret.model_dump_json(indent=2))\n"
},
{
"type": "markdown",
@@ -171,20 +171,20 @@
},
{
"type": "code",
- "source": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-base\"\n\ntry:\n base_model_fs = client.files.filesets.create(\n name=MODEL_NAME,\n description=\"Llama 3.2 1B Instruct base model from HuggingFace\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n token_secret=hf_secret.name,\n ),\n metadata={\n \"model\": {\n \"tool_calling\": {\n \"tool_call_parser\": \"llama3_json\",\n \"auto_tool_choice\": True,\n }\n }\n },\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n\nprint(f\"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}\")\nprint(\"Base model fileset files list:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec.model_dump()}\")",
+ "source": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-base\"\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n description=\"Llama 3.2 1B Instruct base model from HuggingFace\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n token_secret=hf_secret.name,\n ),\n metadata={\n \"model\": {\n \"tool_calling\": {\n \"tool_call_parser\": \"llama3_json\",\n \"auto_tool_choice\": True,\n }\n }\n },\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n\nprint(f\"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}\")\nprint(\"Base model fileset files list:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec.model_dump()}\")",
"language": "python",
- "source_html": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-base"\n\ntry:\n base_model_fs = client.files.filesets.create(\n name=MODEL_NAME,\n description="Llama 3.2 1B Instruct base model from HuggingFace",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n token_secret=hf_secret.name,\n ),\n metadata={\n "model": {\n "tool_calling": {\n "tool_call_parser": "llama3_json",\n "auto_tool_choice": True,\n }\n }\n },\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n\nprint(f"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}")\nprint("Base model fileset files list:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec.model_dump()}")\n"
+ "source_html": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-base"\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n description="Llama 3.2 1B Instruct base model from HuggingFace",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n token_secret=hf_secret.name,\n ),\n metadata={\n "model": {\n "tool_calling": {\n "tool_call_parser": "llama3_json",\n "auto_tool_choice": True,\n }\n }\n },\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n\nprint(f"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}")\nprint("Base model fileset files list:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec.model_dump()}")\n"
},
{
"type": "markdown",
- "source": "## 4. Create LoRA Fine-Tuning Job\n\nCreate a customization job using SFT training with LoRA PEFT. The `peft=LoRaParamsParam()` parameter enables LoRA instead of full-weight fine-tuning.\n\n**LoRA defaults** (can be overridden in `LoRaParamsParam()`):\n- `rank`: LoRA rank (dimensionality of the low-rank matrices)\n- `alpha`: Scaling factor for LoRA updates\n- `target_modules`: Which model layers to apply LoRA to",
- "source_html": "4. Create LoRA Fine-Tuning Job
\nCreate a customization job using SFT training with LoRA PEFT. The peft=LoRaParamsParam() parameter enables LoRA instead of full-weight fine-tuning.
\nLoRA defaults (can be overridden in LoRaParamsParam()):
\n\nrank: LoRA rank (dimensionality of the low-rank matrices) \nalpha: Scaling factor for LoRA updates \ntarget_modules: Which model layers to apply LoRA to \n
\n"
+ "source": "## 4. Create LoRA Fine-Tuning Job\n\nCreate an **Automodel** customization job using `AutomodelJobInput` with `finetuning_type: lora`. After training completes, deploy the base model with LoRA support in the next section.",
+ "source_html": "4. Create LoRA Fine-Tuning Job
\nCreate an Automodel customization job using AutomodelJobInput with finetuning_type: lora. After training completes, deploy the base model with LoRA support in the next section.
\n"
},
{
"type": "code",
- "source": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n DeploymentParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"tool-calling-lora-{job_suffix}\"\n\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n spec=CustomizationJobInputParam(\n model=f\"{WORKSPACE}/{base_model.name}\",\n dataset=f\"fileset://{WORKSPACE}/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=4,\n batch_size=4,\n learning_rate=0.0001,\n max_seq_length=2048,\n micro_batch_size=1,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n ),\n ),\n)\n\nprint(job.model_dump_json(indent=2))",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"tool-calling-lora-{job_suffix}\"\nOUTPUT_NAME = f\"tool-calling-adapter-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"{WORKSPACE}/{base_model.name}\",\n dataset={\"training\": f\"{WORKSPACE}/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 2048,\n },\n schedule={\"epochs\": 4},\n batch={\"global_batch_size\": 4, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": 1e-4},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=WORKSPACE, name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output adapter: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n DeploymentParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"tool-calling-lora-{job_suffix}"\n\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n spec=CustomizationJobInputParam(\n model=f"{WORKSPACE}/{base_model.name}",\n dataset=f"fileset://{WORKSPACE}/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=4,\n batch_size=4,\n learning_rate=0.0001,\n max_seq_length=2048,\n micro_batch_size=1,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n ),\n ),\n)\n\nprint(job.model_dump_json(indent=2))\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"tool-calling-lora-{job_suffix}"\nOUTPUT_NAME = f"tool-calling-adapter-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"{WORKSPACE}/{base_model.name}",\n dataset={"training": f"{WORKSPACE}/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 2048,\n },\n schedule={"epochs": 4},\n batch={"global_batch_size": 4, "micro_batch_size": 1},\n optimizer={"learning_rate": 1e-4},\n parallelism={"num_gpus_per_node": 1},\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=WORKSPACE, name=JOB_NAME\n)\n\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output adapter: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -193,26 +193,26 @@
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {\"completed\", \"cancelled\", \"error\"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n \"\"\"Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n \"\"\"\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f\"[{label}] Job: {status.name}\")\n print(f\"[{label}] Status: {status.status}\")\n print(f\"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s\")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f\"\\n[{label}] Job finished: {status.status}\")\n return status\n\n if elapsed > timeout:\n print(f\"\\n[{label}] Timeout after {timeout_minutes} minutes\")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n \"\"\"Extract and display training step progress.\"\"\"\n for step in status.steps or []:\n if step.name == \"customization-training-job\":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get(\"step\"), details.get(\"max_steps\")\n if s is not None and mx is not None:\n print(f\"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)\")\n if phase := details.get(\"phase\"):\n print(f\"Phase: {phase}\")\n return\n print(\"Training step not started yet\")\n\nprint(\"Defined wait_for_job helper function\")",
+ "source": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {\"completed\", \"failed\", \"cancelled\", \"error\"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n \"\"\"Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n \"\"\"\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f\"[{label}] Job: {status.name}\")\n print(f\"[{label}] Status: {status.status}\")\n print(f\"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s\")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f\"\\n[{label}] Job finished: {status.status}\")\n return status\n\n if elapsed > timeout:\n print(f\"\\n[{label}] Timeout after {timeout_minutes} minutes\")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n \"\"\"Extract and display training step progress.\"\"\"\n for step in status.steps or []:\n if step.name == \"training\":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get(\"step\"), details.get(\"max_steps\")\n if s is not None and mx is not None:\n print(f\"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)\")\n if phase := details.get(\"phase\"):\n print(f\"Phase: {phase}\")\n return\n print(\"Training step not started yet\")\n\nprint(\"Defined wait_for_job helper function\")",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {"completed", "cancelled", "error"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n """Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n """\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f"[{label}] Job: {status.name}")\n print(f"[{label}] Status: {status.status}")\n print(f"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f"\\n[{label}] Job finished: {status.status}")\n return status\n\n if elapsed > timeout:\n print(f"\\n[{label}] Timeout after {timeout_minutes} minutes")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n """Extract and display training step progress."""\n for step in status.steps or []:\n if step.name == "customization-training-job":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get("step"), details.get("max_steps")\n if s is not None and mx is not None:\n print(f"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)")\n if phase := details.get("phase"):\n print(f"Phase: {phase}")\n return\n print("Training step not started yet")\n\nprint("Defined wait_for_job helper function")\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {"completed", "failed", "cancelled", "error"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n """Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n """\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f"[{label}] Job: {status.name}")\n print(f"[{label}] Status: {status.status}")\n print(f"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f"\\n[{label}] Job finished: {status.status}")\n return status\n\n if elapsed > timeout:\n print(f"\\n[{label}] Timeout after {timeout_minutes} minutes")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n """Extract and display training step progress."""\n for step in status.steps or []:\n if step.name == "training":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get("step"), details.get("max_steps")\n if s is not None and mx is not None:\n print(f"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)")\n if phase := details.get("phase"):\n print(f"Phase: {phase}")\n return\n print("Training step not started yet")\n\nprint("Defined wait_for_job helper function")\n"
},
{
"type": "code",
- "source": "job_status = wait_for_job(\n poll_fn=lambda: client.customization.jobs.get_status(name=job.name),\n label=\"Training\",\n timeout_minutes=120,\n display_fn=training_progress,\n)",
+ "source": "job_status = wait_for_job(\n poll_fn=lambda: client.jobs.get_status(name=job.job.name, workspace=WORKSPACE),\n label=\"Training\",\n timeout_minutes=120,\n display_fn=training_progress,\n)\n\nif job_status.status != \"completed\":\n raise RuntimeError(f\"Training job finished with status: {job_status.status}\")",
"language": "python",
- "source_html": "job_status = wait_for_job(\n poll_fn=lambda: client.customization.jobs.get_status(name=job.name),\n label="Training",\n timeout_minutes=120,\n display_fn=training_progress,\n)\n"
+ "source_html": "job_status = wait_for_job(\n poll_fn=lambda: client.jobs.get_status(name=job.job.name, workspace=WORKSPACE),\n label="Training",\n timeout_minutes=120,\n display_fn=training_progress,\n)\n\nif job_status.status != "completed":\n raise RuntimeError(f"Training job finished with status: {job_status.status}")\n"
},
{
"type": "markdown",
- "source": "## 5. Verify Auto-Deployed Model\n\nSince we set `lora_enabled=True` in the customization job's `deployment_config`, the platform automatically creates a NIM deployment for the base model after training completes. The LoRA adapter is attached to the base model entity (enabled by default) and the deployment serves both the base weights and the adapter through a single NIM instance.",
- "source_html": "5. Verify Auto-Deployed Model
\nSince we set lora_enabled=True in the customization job's deployment_config, the platform automatically creates a NIM deployment for the base model after training completes. The LoRA adapter is attached to the base model entity (enabled by default) and the deployment serves both the base weights and the adapter through a single NIM instance.
\n"
+ "source": "## 5. Deploy Fine-Tuned Model\n\nAfter training completes, verify the LoRA adapter is attached to the base model entity, then create a NIM deployment with `lora_enabled=True` so both the base weights and adapter are served through a single deployment.",
+ "source_html": "5. Deploy Fine-Tuned Model
\nAfter training completes, verify the LoRA adapter is attached to the base model entity, then create a NIM deployment with lora_enabled=True so both the base weights and adapter are served through a single deployment.
\n"
},
{
"type": "code",
- "source": "ADAPTER_NAME = job.spec.output.name\nprint(f\"Looking for adapter: {ADAPTER_NAME}\")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f\"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)\")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f\"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s\"\n )\n\nprint(f\"Base model: {base_model.name}\")\nprint(f\"Adapter:\\n{adapter.model_dump_json(indent=2)}\")",
+ "source": "ADAPTER_NAME = OUTPUT_NAME\nprint(f\"Looking for adapter: {ADAPTER_NAME}\")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME, workspace=WORKSPACE)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f\"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)\")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f\"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s\"\n )\n\nprint(f\"Base model: {base_model.name}\")\nprint(f\"Adapter:\\n{adapter.model_dump_json(indent=2)}\")\n\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"tool-calling-deploy-cfg-{deploy_suffix}\"\ndeployment_name = f\"tool-calling-deploy-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=WORKSPACE,\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": WORKSPACE,\n \"model_name\": MODEL_NAME,\n \"lora_enabled\": True,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n \"additional_args\": [\"--max-lora-rank\", \"32\"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace=WORKSPACE,\n name=deployment_name,\n config=deployment_config.name,\n)\n\nprint(f\"Deployment status: {deployment.status}\")",
"language": "python",
- "source_html": "ADAPTER_NAME = job.spec.output.name\nprint(f"Looking for adapter: {ADAPTER_NAME}")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s"\n )\n\nprint(f"Base model: {base_model.name}")\nprint(f"Adapter:\\n{adapter.model_dump_json(indent=2)}")\n"
+ "source_html": "ADAPTER_NAME = OUTPUT_NAME\nprint(f"Looking for adapter: {ADAPTER_NAME}")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME, workspace=WORKSPACE)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s"\n )\n\nprint(f"Base model: {base_model.name}")\nprint(f"Adapter:\\n{adapter.model_dump_json(indent=2)}")\n\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"tool-calling-deploy-cfg-{deploy_suffix}"\ndeployment_name = f"tool-calling-deploy-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=WORKSPACE,\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={\n "model_namespace": WORKSPACE,\n "model_name": MODEL_NAME,\n "lora_enabled": True,\n },\n executor_config={\n "gpu": 1,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n "additional_args": ["--max-lora-rank", "32"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace=WORKSPACE,\n name=deployment_name,\n config=deployment_config.name,\n)\n\nprint(f"Deployment status: {deployment.status}")\n"
},
{
"type": "markdown",
@@ -221,9 +221,9 @@
},
{
"type": "code",
- "source": "DEPLOYMENT_NAME = f\"sft-deploy-{MODEL_NAME}\"\n\nTIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f\"Monitoring deployment '{DEPLOYMENT_NAME}'...\")\nprint(f\"Timeout: {TIMEOUT_MINUTES} minutes\\n\")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=DEPLOYMENT_NAME,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f\"Deployment: {DEPLOYMENT_NAME}\")\n print(f\"Status: {deployment_status.status}\")\n print(f\"Elapsed time: {elapsed_min}m {elapsed_sec}s\")\n\n if deployment_status.status == \"READY\":\n print(\"\\nDeployment is ready!\")\n if not client.models.wait_for_gateway(DEPLOYMENT_NAME, workspace=WORKSPACE, timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n\n if deployment_status.status in (\"FAILED\", \"ERROR\", \"TERMINATED\", \"LOST\", \"DELETED\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f\"Deployment timeout after {TIMEOUT_MINUTES} minutes\")\n\n time.sleep(15)",
+ "source": "TIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f\"Monitoring deployment '{deployment_name}'...\")\nprint(f\"Timeout: {TIMEOUT_MINUTES} minutes\\n\")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=deployment_name,\n workspace=WORKSPACE,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f\"Deployment: {deployment_name}\")\n print(f\"Status: {deployment_status.status}\")\n print(f\"Elapsed time: {elapsed_min}m {elapsed_sec}s\")\n\n if deployment_status.status in (\"RUNNING\", \"READY\"):\n print(\"\\nDeployment is ready!\")\n if not client.models.wait_for_gateway(deployment_name, workspace=WORKSPACE, timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n\n if deployment_status.status in (\"FAILED\", \"ERROR\", \"TERMINATED\", \"LOST\", \"DELETED\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f\"Deployment timeout after {TIMEOUT_MINUTES} minutes\")\n\n time.sleep(15)",
"language": "python",
- "source_html": "DEPLOYMENT_NAME = f"sft-deploy-{MODEL_NAME}"\n\nTIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f"Monitoring deployment '{DEPLOYMENT_NAME}'...")\nprint(f"Timeout: {TIMEOUT_MINUTES} minutes\\n")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=DEPLOYMENT_NAME,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f"Deployment: {DEPLOYMENT_NAME}")\n print(f"Status: {deployment_status.status}")\n print(f"Elapsed time: {elapsed_min}m {elapsed_sec}s")\n\n if deployment_status.status == "READY":\n print("\\nDeployment is ready!")\n if not client.models.wait_for_gateway(DEPLOYMENT_NAME, workspace=WORKSPACE, timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n\n if deployment_status.status in ("FAILED", "ERROR", "TERMINATED", "LOST", "DELETED"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f"Deployment timeout after {TIMEOUT_MINUTES} minutes")\n\n time.sleep(15)\n"
+ "source_html": "TIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f"Monitoring deployment '{deployment_name}'...")\nprint(f"Timeout: {TIMEOUT_MINUTES} minutes\\n")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=deployment_name,\n workspace=WORKSPACE,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f"Deployment: {deployment_name}")\n print(f"Status: {deployment_status.status}")\n print(f"Elapsed time: {elapsed_min}m {elapsed_sec}s")\n\n if deployment_status.status in ("RUNNING", "READY"):\n print("\\nDeployment is ready!")\n if not client.models.wait_for_gateway(deployment_name, workspace=WORKSPACE, timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n\n if deployment_status.status in ("FAILED", "ERROR", "TERMINATED", "LOST", "DELETED"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f"Deployment timeout after {TIMEOUT_MINUTES} minutes")\n\n time.sleep(15)\n"
},
{
"type": "markdown",
@@ -232,9 +232,9 @@
},
{
"type": "code",
- "source": "test_messages = [\n {\"role\": \"user\", \"content\": \"Calculate the factorial of 12 using math functions.\"},\n]\n\ntest_tools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"math_factorial\",\n \"description\": \"Calculate the factorial of a given number.\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"number\": {\n \"type\": \"integer\",\n \"description\": \"The number for which factorial needs to be calculated.\",\n }\n },\n \"required\": [\"number\"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n \"\"\"Send a tool calling request and display the response.\"\"\"\n response = client.inference.gateway.model.post(\n \"v1/chat/completions\",\n name=model_name,\n body={\n \"messages\": test_messages,\n \"tools\": test_tools,\n \"tool_choice\": \"auto\",\n \"temperature\": 0,\n \"max_tokens\": 256,\n },\n )\n\n print(f\"{'=' * 60}\")\n print(f\" {label}\")\n print(f\"{'=' * 60}\")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, \"BASE MODEL (before fine-tuning)\")\ntest_tool_calling(ADAPTER_NAME, \"FINE-TUNED MODEL (after LoRA)\")",
+ "source": "test_messages = [\n {\"role\": \"user\", \"content\": \"Calculate the factorial of 12 using math functions.\"},\n]\n\ntest_tools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"math_factorial\",\n \"description\": \"Calculate the factorial of a given number.\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"number\": {\n \"type\": \"integer\",\n \"description\": \"The number for which factorial needs to be calculated.\",\n }\n },\n \"required\": [\"number\"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n \"\"\"Send a tool calling request and display the response.\"\"\"\n response = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=WORKSPACE,\n body={\n \"model\": model_name,\n \"messages\": test_messages,\n \"tools\": test_tools,\n \"tool_choice\": \"auto\",\n \"temperature\": 0,\n \"max_tokens\": 256,\n },\n )\n\n print(f\"{'=' * 60}\")\n print(f\" {label}\")\n print(f\"{'=' * 60}\")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, \"BASE MODEL (before fine-tuning)\")\ntest_tool_calling(ADAPTER_NAME, \"FINE-TUNED MODEL (after LoRA)\")",
"language": "python",
- "source_html": "test_messages = [\n {"role": "user", "content": "Calculate the factorial of 12 using math functions."},\n]\n\ntest_tools = [\n {\n "type": "function",\n "function": {\n "name": "math_factorial",\n "description": "Calculate the factorial of a given number.",\n "parameters": {\n "type": "object",\n "properties": {\n "number": {\n "type": "integer",\n "description": "The number for which factorial needs to be calculated.",\n }\n },\n "required": ["number"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n """Send a tool calling request and display the response."""\n response = client.inference.gateway.model.post(\n "v1/chat/completions",\n name=model_name,\n body={\n "messages": test_messages,\n "tools": test_tools,\n "tool_choice": "auto",\n "temperature": 0,\n "max_tokens": 256,\n },\n )\n\n print(f"{'=' * 60}")\n print(f" {label}")\n print(f"{'=' * 60}")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, "BASE MODEL (before fine-tuning)")\ntest_tool_calling(ADAPTER_NAME, "FINE-TUNED MODEL (after LoRA)")\n"
+ "source_html": "test_messages = [\n {"role": "user", "content": "Calculate the factorial of 12 using math functions."},\n]\n\ntest_tools = [\n {\n "type": "function",\n "function": {\n "name": "math_factorial",\n "description": "Calculate the factorial of a given number.",\n "parameters": {\n "type": "object",\n "properties": {\n "number": {\n "type": "integer",\n "description": "The number for which factorial needs to be calculated.",\n }\n },\n "required": ["number"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n """Send a tool calling request and display the response."""\n response = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace=WORKSPACE,\n body={\n "model": model_name,\n "messages": test_messages,\n "tools": test_tools,\n "tool_choice": "auto",\n "temperature": 0,\n "max_tokens": 256,\n },\n )\n\n print(f"{'=' * 60}")\n print(f" {label}")\n print(f"{'=' * 60}")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, "BASE MODEL (before fine-tuning)")\ntest_tool_calling(ADAPTER_NAME, "FINE-TUNED MODEL (after LoRA)")\n"
},
{
"type": "markdown",
diff --git a/docs/fern/components/notebooks/tool-calling.ts b/docs/fern/components/notebooks/tool-calling.ts
index 65d63f5e36..70fb87d896 100644
--- a/docs/fern/components/notebooks/tool-calling.ts
+++ b/docs/fern/components/notebooks/tool-calling.ts
@@ -1,7 +1,9 @@
-// SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
-// SPDX-License-Identifier: Apache-2.0
-
-/** Auto-generated by ipynb-to-fern-json.py - do not edit */
+/**
+ * SPDX-FileCopyrightText: Copyright (c) 2025-2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+ * SPDX-License-Identifier: Apache-2.0
+ *
+ * Auto-generated by ipynb-to-fern-json.py - do not edit manually.
+ */
export default { cells: [
{
"type": "markdown",
@@ -10,8 +12,8 @@ export default { cells: [
},
{
"type": "markdown",
- "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Installed the Python SDK with the Data Designer extra** (`uv pip install nemo-platform[data-designer]`)\n1. **Completed the [Quickstart](../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n1. (Optional if running outside of Quickstart) **Authenticated with the platform** using the CLI:\n ```bash\n nemo auth login\n ```\n For non-default URLs: `nemo auth login --base-url `\n1. Set your **Hugging Face token** as an environment variable before running the notebook (`export HF_TOKEN=\"\"`)\n1. **Accepted the dataset license** for [Salesforce/xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k) on Hugging Face (used for evaluation) ",
- "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Installed the Python SDK with the Data Designer extra (
uv pip install nemo-platform[data-designer]) \n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- (Optional if running outside of Quickstart) Authenticated with the platform using the CLI:
nemo auth login\n
\nFor non-default URLs: nemo auth login --base-url <YOUR_NMP_BASE_URL> \n- Set your Hugging Face token as an environment variable before running the notebook (
export HF_TOKEN="<your-token>") \n- Accepted the dataset license for Salesforce/xlam-function-calling-60k on Hugging Face (used for evaluation)
\n
\n"
+ "source": "## Prerequisites\n\nBefore starting this tutorial, ensure you have:\n\n1. **Installed the Python SDK with Data Designer support** (PyPI wrapper: `uv pip install \"nemo-platform[all]\"`; source checkout: run `make bootstrap` from the repository root)\n1. **Completed the [Quickstart](../get-started/quickstart.md)** to install and deploy NeMo Platform locally\n1. (Optional if running outside of Quickstart) **Authenticated with the platform** using the CLI:\n ```bash\n nemo auth login\n ```\n For non-default URLs: `nemo auth login --base-url `\n1. Set your **Hugging Face token** as an environment variable before running the notebook (`export HF_TOKEN=\"\"`)\n1. **Accepted the dataset license** for [Salesforce/xlam-function-calling-60k](https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k) on Hugging Face (used for evaluation) ",
+ "source_html": "Prerequisites
\nBefore starting this tutorial, ensure you have:
\n\n- Installed the Python SDK with Data Designer support (PyPI wrapper:
uv pip install "nemo-platform[all]"; source checkout: run make bootstrap from the repository root) \n- Completed the Quickstart to install and deploy NeMo Platform locally
\n- (Optional if running outside of Quickstart) Authenticated with the platform using the CLI:
nemo auth login\n
\nFor non-default URLs: nemo auth login --base-url <YOUR_NMP_BASE_URL> \n- Set your Hugging Face token as an environment variable before running the notebook (
export HF_TOKEN="<your-token>") \n- Accepted the dataset license for Salesforce/xlam-function-calling-60k on Hugging Face (used for evaluation)
\n
\n"
},
{
"type": "markdown",
@@ -141,9 +143,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "DATASET_NAME = \"tool-calling-dataset\"\n\ntry:\n client.files.filesets.create(\n name=DATASET_NAME,\n description=\"synthetic tool-calling training and validation data in OpenAI chat format\",\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path=\"\",\n fileset=DATASET_NAME,\n)\n\nprint(\"\\nTraining data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME).data],\n indent=2,\n))",
+ "source": "DATASET_NAME = \"tool-calling-dataset\"\n\ntry:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=DATASET_NAME,\n description=\"synthetic tool-calling training and validation data in OpenAI chat format\",\n )\n print(f\"Created fileset: {DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path=\"\",\n fileset=DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint(\"\\nTraining data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))",
"language": "python",
- "source_html": "DATASET_NAME = "tool-calling-dataset"\n\ntry:\n client.files.filesets.create(\n name=DATASET_NAME,\n description="synthetic tool-calling training and validation data in OpenAI chat format",\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path="",\n fileset=DATASET_NAME,\n)\n\nprint("\\nTraining data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME).data],\n indent=2,\n))\n"
+ "source_html": "DATASET_NAME = "tool-calling-dataset"\n\ntry:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=DATASET_NAME,\n description="synthetic tool-calling training and validation data in OpenAI chat format",\n )\n print(f"Created fileset: {DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=DATASET_PATH,\n remote_path="",\n fileset=DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint("\\nTraining data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n"
},
{
"type": "markdown",
@@ -152,9 +154,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "try:\n client.files.filesets.create(\n name=EVAL_DATASET_NAME,\n description=\"synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)\",\n )\n print(f\"Created fileset: {EVAL_DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path=\"\",\n fileset=EVAL_DATASET_NAME,\n)\n\nprint(\"\\nEvaluation data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME).data],\n indent=2,\n))",
+ "source": "try:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=EVAL_DATASET_NAME,\n description=\"synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)\",\n )\n print(f\"Created fileset: {EVAL_DATASET_NAME}\")\nexcept ConflictError:\n print(f\"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...\")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path=\"\",\n fileset=EVAL_DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint(\"\\nEvaluation data files:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))",
"language": "python",
- "source_html": "try:\n client.files.filesets.create(\n name=EVAL_DATASET_NAME,\n description="synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)",\n )\n print(f"Created fileset: {EVAL_DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path="",\n fileset=EVAL_DATASET_NAME,\n)\n\nprint("\\nEvaluation data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME).data],\n indent=2,\n))\n"
+ "source_html": "try:\n client.files.filesets.create(\n workspace=WORKSPACE,\n name=EVAL_DATASET_NAME,\n description="synthetic tool-calling evaluation data (messages, tools, ground-truth tool_calls)",\n )\n print(f"Created fileset: {EVAL_DATASET_NAME}")\nexcept ConflictError:\n print(f"Fileset '{EVAL_DATASET_NAME}' already exists, continuing...")\n\nclient.files.upload(\n local_path=EVAL_DATASET_PATH,\n remote_path="",\n fileset=EVAL_DATASET_NAME,\n workspace=WORKSPACE,\n)\n\nprint("\\nEvaluation data files:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=EVAL_DATASET_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n"
},
{
"type": "markdown",
@@ -163,9 +165,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "HF_TOKEN = os.environ.get(\"HF_TOKEN\")\nif HF_TOKEN is None:\n raise ValueError(\"HF_TOKEN is not set.\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n name=name,\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name)\n\n\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(hf_secret.model_dump_json(indent=2))",
+ "source": "HF_TOKEN = os.environ.get(\"HF_TOKEN\")\nif HF_TOKEN is None:\n raise ValueError(\"HF_TOKEN is not set.\")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f\"{label} is not set\")\n try:\n secret = client.secrets.create(\n workspace=WORKSPACE,\n name=name,\n value=value,\n )\n print(f\"Created secret: {name}\")\n return secret\n except ConflictError:\n print(f\"Secret '{name}' already exists, continuing...\")\n return client.secrets.retrieve(name=name, workspace=WORKSPACE)\n\n\nhf_secret = create_or_get_secret(\"hf-token\", HF_TOKEN, \"HF_TOKEN\")\nprint(hf_secret.model_dump_json(indent=2))",
"language": "python",
- "source_html": "HF_TOKEN = os.environ.get("HF_TOKEN")\nif HF_TOKEN is None:\n raise ValueError("HF_TOKEN is not set.")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n name=name,\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name)\n\n\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(hf_secret.model_dump_json(indent=2))\n"
+ "source_html": "HF_TOKEN = os.environ.get("HF_TOKEN")\nif HF_TOKEN is None:\n raise ValueError("HF_TOKEN is not set.")\n\n\ndef create_or_get_secret(name: str, value: str | None, label: str):\n if not value:\n raise ValueError(f"{label} is not set")\n try:\n secret = client.secrets.create(\n workspace=WORKSPACE,\n name=name,\n value=value,\n )\n print(f"Created secret: {name}")\n return secret\n except ConflictError:\n print(f"Secret '{name}' already exists, continuing...")\n return client.secrets.retrieve(name=name, workspace=WORKSPACE)\n\n\nhf_secret = create_or_get_secret("hf-token", HF_TOKEN, "HF_TOKEN")\nprint(hf_secret.model_dump_json(indent=2))\n"
},
{
"type": "markdown",
@@ -174,20 +176,20 @@ export default { cells: [
},
{
"type": "code",
- "source": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-base\"\n\ntry:\n base_model_fs = client.files.filesets.create(\n name=MODEL_NAME,\n description=\"Llama 3.2 1B Instruct base model from HuggingFace\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n token_secret=hf_secret.name,\n ),\n metadata={\n \"model\": {\n \"tool_calling\": {\n \"tool_call_parser\": \"llama3_json\",\n \"auto_tool_choice\": True,\n }\n }\n },\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n\nprint(f\"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}\")\nprint(\"Base model fileset files list:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec.model_dump()}\")",
+ "source": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = \"meta-llama/Llama-3.2-1B-Instruct\"\nMODEL_NAME = \"llama-3-2-1b-base\"\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n description=\"Llama 3.2 1B Instruct base model from HuggingFace\",\n storage=HuggingfaceStorageConfigParam(\n type=\"huggingface\",\n repo_id=HF_REPO_ID,\n repo_type=\"model\",\n token_secret=hf_secret.name,\n ),\n metadata={\n \"model\": {\n \"tool_calling\": {\n \"tool_call_parser\": \"llama3_json\",\n \"auto_tool_choice\": True,\n }\n }\n },\n )\n print(f\"Created base model fileset: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model fileset already exists. Skipping creation.\")\n base_model_fs = client.files.filesets.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n print(f\"Created Model Entity: {MODEL_NAME}\")\nexcept ConflictError:\n print(f\"Base model already exists. Updating fileset if different.\")\n base_model = client.models.update(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f\"{WORKSPACE}/{MODEL_NAME}\",\n )\n\nprint(f\"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}\")\nprint(\"Base model fileset files list:\")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint(\"\\nWaiting for ModelSpec to be populated...\")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f\"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds\")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\nprint(f\"ModelSpec populated: {base_model.spec.model_dump()}\")",
"language": "python",
- "source_html": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-base"\n\ntry:\n base_model_fs = client.files.filesets.create(\n name=MODEL_NAME,\n description="Llama 3.2 1B Instruct base model from HuggingFace",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n token_secret=hf_secret.name,\n ),\n metadata={\n "model": {\n "tool_calling": {\n "tool_call_parser": "llama3_json",\n "auto_tool_choice": True,\n }\n }\n },\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n\nprint(f"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}")\nprint("Base model fileset files list:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec.model_dump()}")\n"
+ "source_html": "import time\n\nfrom nemo_platform.types.files import HuggingfaceStorageConfigParam\n\nHF_REPO_ID = "meta-llama/Llama-3.2-1B-Instruct"\nMODEL_NAME = "llama-3-2-1b-base"\n\ntry:\n base_model_fs = client.files.filesets.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n description="Llama 3.2 1B Instruct base model from HuggingFace",\n storage=HuggingfaceStorageConfigParam(\n type="huggingface",\n repo_id=HF_REPO_ID,\n repo_type="model",\n token_secret=hf_secret.name,\n ),\n metadata={\n "model": {\n "tool_calling": {\n "tool_call_parser": "llama3_json",\n "auto_tool_choice": True,\n }\n }\n },\n )\n print(f"Created base model fileset: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model fileset already exists. Skipping creation.")\n base_model_fs = client.files.filesets.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\ntry:\n base_model = client.models.create(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n print(f"Created Model Entity: {MODEL_NAME}")\nexcept ConflictError:\n print(f"Base model already exists. Updating fileset if different.")\n base_model = client.models.update(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n fileset=f"{WORKSPACE}/{MODEL_NAME}",\n )\n\nprint(f"\\nBase model fileset: fileset://{WORKSPACE}/{base_model.name}")\nprint("Base model fileset files list:")\nprint(json.dumps(\n [f.model_dump() for f in client.files.list(fileset=MODEL_NAME, workspace=WORKSPACE).data],\n indent=2,\n))\n\n# Wait for ModelSpec to be populated from the checkpoint\nprint("\\nWaiting for ModelSpec to be populated...")\nSPEC_TIMEOUT_SECONDS = 120\nspec_start = time.time()\nwhile not base_model.spec:\n if time.time() - spec_start > SPEC_TIMEOUT_SECONDS:\n raise TimeoutError(f"ModelSpec not populated within {SPEC_TIMEOUT_SECONDS} seconds")\n time.sleep(2)\n base_model = client.models.retrieve(\n workspace=WORKSPACE,\n name=MODEL_NAME,\n )\n\nprint(f"ModelSpec populated: {base_model.spec.model_dump()}")\n"
},
{
"type": "markdown",
- "source": "## 4. Create LoRA Fine-Tuning Job\n\nCreate a customization job using SFT training with LoRA PEFT. The `peft=LoRaParamsParam()` parameter enables LoRA instead of full-weight fine-tuning.\n\n**LoRA defaults** (can be overridden in `LoRaParamsParam()`):\n- `rank`: LoRA rank (dimensionality of the low-rank matrices)\n- `alpha`: Scaling factor for LoRA updates\n- `target_modules`: Which model layers to apply LoRA to",
- "source_html": "4. Create LoRA Fine-Tuning Job
\nCreate a customization job using SFT training with LoRA PEFT. The peft=LoRaParamsParam() parameter enables LoRA instead of full-weight fine-tuning.
\nLoRA defaults (can be overridden in LoRaParamsParam()):
\n\nrank: LoRA rank (dimensionality of the low-rank matrices) \nalpha: Scaling factor for LoRA updates \ntarget_modules: Which model layers to apply LoRA to \n
\n"
+ "source": "## 4. Create LoRA Fine-Tuning Job\n\nCreate an **Automodel** customization job using `AutomodelJobInput` with `finetuning_type: lora`. After training completes, deploy the base model with LoRA support in the next section.",
+ "source_html": "4. Create LoRA Fine-Tuning Job
\nCreate an Automodel customization job using AutomodelJobInput with finetuning_type: lora. After training completes, deploy the base model with LoRA support in the next section.
\n"
},
{
"type": "code",
- "source": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n DeploymentParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"tool-calling-lora-{job_suffix}\"\n\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n spec=CustomizationJobInputParam(\n model=f\"{WORKSPACE}/{base_model.name}\",\n dataset=f\"fileset://{WORKSPACE}/{DATASET_NAME}\",\n training=SftTrainingParam(\n type=\"sft\",\n epochs=4,\n batch_size=4,\n learning_rate=0.0001,\n max_seq_length=2048,\n micro_batch_size=1,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n ),\n ),\n)\n\nprint(job.model_dump_json(indent=2))",
+ "source": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f\"tool-calling-lora-{job_suffix}\"\nOUTPUT_NAME = f\"tool-calling-adapter-{job_suffix}\"\n\nspec = AutomodelJobInput(\n model=f\"{WORKSPACE}/{base_model.name}\",\n dataset={\"training\": f\"{WORKSPACE}/{DATASET_NAME}\"},\n training={\n \"training_type\": \"sft\",\n \"finetuning_type\": \"lora\",\n \"max_seq_length\": 2048,\n },\n schedule={\"epochs\": 4},\n batch={\"global_batch_size\": 4, \"micro_batch_size\": 1},\n optimizer={\"learning_rate\": 1e-4},\n parallelism={\"num_gpus_per_node\": 1},\n output={\"name\": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=WORKSPACE, name=JOB_NAME\n)\n\nprint(f\"Submitted job: {job.job.name}\")\nprint(f\"Output adapter: {OUTPUT_NAME}\")",
"language": "python",
- "source_html": "import uuid\nfrom nemo_platform.types.customization import (\n CustomizationJobInputParam,\n SftTrainingParam,\n ParallelismParamsParam,\n LoRaParamsParam,\n DeploymentParamsParam,\n)\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"tool-calling-lora-{job_suffix}"\n\n\njob = client.customization.jobs.create(\n name=JOB_NAME,\n spec=CustomizationJobInputParam(\n model=f"{WORKSPACE}/{base_model.name}",\n dataset=f"fileset://{WORKSPACE}/{DATASET_NAME}",\n training=SftTrainingParam(\n type="sft",\n epochs=4,\n batch_size=4,\n learning_rate=0.0001,\n max_seq_length=2048,\n micro_batch_size=1,\n peft=LoRaParamsParam(),\n parallelism=ParallelismParamsParam(\n num_gpus_per_node=1,\n num_nodes=1,\n tensor_parallel_size=1,\n pipeline_parallel_size=1,\n ),\n ),\n deployment_config=DeploymentParamsParam(\n lora_enabled=True,\n ),\n ),\n)\n\nprint(job.model_dump_json(indent=2))\n"
+ "source_html": "import uuid\nfrom nemo_automodel_plugin.schema import AutomodelJobInput\n\njob_suffix = uuid.uuid4().hex[:4]\nJOB_NAME = f"tool-calling-lora-{job_suffix}"\nOUTPUT_NAME = f"tool-calling-adapter-{job_suffix}"\n\nspec = AutomodelJobInput(\n model=f"{WORKSPACE}/{base_model.name}",\n dataset={"training": f"{WORKSPACE}/{DATASET_NAME}"},\n training={\n "training_type": "sft",\n "finetuning_type": "lora",\n "max_seq_length": 2048,\n },\n schedule={"epochs": 4},\n batch={"global_batch_size": 4, "micro_batch_size": 1},\n optimizer={"learning_rate": 1e-4},\n parallelism={"num_gpus_per_node": 1},\n output={"name": OUTPUT_NAME},\n)\n\njob = client.customization.automodel.jobs.create(\n spec=spec, workspace=WORKSPACE, name=JOB_NAME\n)\n\nprint(f"Submitted job: {job.job.name}")\nprint(f"Output adapter: {OUTPUT_NAME}")\n"
},
{
"type": "markdown",
@@ -196,26 +198,26 @@ export default { cells: [
},
{
"type": "code",
- "source": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {\"completed\", \"cancelled\", \"error\"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n \"\"\"Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n \"\"\"\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f\"[{label}] Job: {status.name}\")\n print(f\"[{label}] Status: {status.status}\")\n print(f\"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s\")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f\"\\n[{label}] Job finished: {status.status}\")\n return status\n\n if elapsed > timeout:\n print(f\"\\n[{label}] Timeout after {timeout_minutes} minutes\")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n \"\"\"Extract and display training step progress.\"\"\"\n for step in status.steps or []:\n if step.name == \"customization-training-job\":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get(\"step\"), details.get(\"max_steps\")\n if s is not None and mx is not None:\n print(f\"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)\")\n if phase := details.get(\"phase\"):\n print(f\"Phase: {phase}\")\n return\n print(\"Training step not started yet\")\n\nprint(\"Defined wait_for_job helper function\")",
+ "source": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {\"completed\", \"failed\", \"cancelled\", \"error\"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n \"\"\"Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n \"\"\"\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f\"[{label}] Job: {status.name}\")\n print(f\"[{label}] Status: {status.status}\")\n print(f\"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s\")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f\"\\n[{label}] Job finished: {status.status}\")\n return status\n\n if elapsed > timeout:\n print(f\"\\n[{label}] Timeout after {timeout_minutes} minutes\")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n \"\"\"Extract and display training step progress.\"\"\"\n for step in status.steps or []:\n if step.name == \"training\":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get(\"step\"), details.get(\"max_steps\")\n if s is not None and mx is not None:\n print(f\"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)\")\n if phase := details.get(\"phase\"):\n print(f\"Phase: {phase}\")\n return\n print(\"Training step not started yet\")\n\nprint(\"Defined wait_for_job helper function\")",
"language": "python",
- "source_html": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {"completed", "cancelled", "error"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n """Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n """\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f"[{label}] Job: {status.name}")\n print(f"[{label}] Status: {status.status}")\n print(f"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f"\\n[{label}] Job finished: {status.status}")\n return status\n\n if elapsed > timeout:\n print(f"\\n[{label}] Timeout after {timeout_minutes} minutes")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n """Extract and display training step progress."""\n for step in status.steps or []:\n if step.name == "customization-training-job":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get("step"), details.get("max_steps")\n if s is not None and mx is not None:\n print(f"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)")\n if phase := details.get("phase"):\n print(f"Phase: {phase}")\n return\n print("Training step not started yet")\n\nprint("Defined wait_for_job helper function")\n"
+ "source_html": "import time\nfrom IPython.display import clear_output\n\nTERMINAL_JOB_STATUSES = {"completed", "failed", "cancelled", "error"}\n\n\ndef wait_for_job(poll_fn, label, timeout_minutes=60, poll_interval=10, display_fn=None):\n """Poll a platform job until it reaches a terminal state.\n\n Both Customizer and Evaluator use the Core Jobs service, so the same\n PlatformJobStatus values apply: created, pending, active, cancelled,\n cancelling, error, completed, paused, pausing, resuming.\n\n Args:\n poll_fn: Callable returning a status object with .name and .status attributes.\n label: Display label for progress output.\n timeout_minutes: Maximum time to wait before returning.\n poll_interval: Seconds between polls.\n display_fn: Optional callable(status) to print extra details after the header.\n\n Returns:\n The final status object.\n """\n start = time.time()\n timeout = timeout_minutes * 60\n\n while True:\n status = poll_fn()\n elapsed = time.time() - start\n elapsed_min, elapsed_sec = divmod(int(elapsed), 60)\n\n clear_output(wait=True)\n print(f"[{label}] Job: {status.name}")\n print(f"[{label}] Status: {status.status}")\n print(f"[{label}] Elapsed: {elapsed_min}m {elapsed_sec}s")\n\n if display_fn:\n display_fn(status)\n\n if status.status in TERMINAL_JOB_STATUSES:\n print(f"\\n[{label}] Job finished: {status.status}")\n return status\n\n if elapsed > timeout:\n print(f"\\n[{label}] Timeout after {timeout_minutes} minutes")\n return status\n\n time.sleep(poll_interval)\n\n\ndef training_progress(status):\n """Extract and display training step progress."""\n for step in status.steps or []:\n if step.name == "training":\n for task in step.tasks or []:\n details = task.status_details or {}\n s, mx = details.get("step"), details.get("max_steps")\n if s is not None and mx is not None:\n print(f"Training: Step {s}/{mx} ({(s / mx) * 100:.1f}%)")\n if phase := details.get("phase"):\n print(f"Phase: {phase}")\n return\n print("Training step not started yet")\n\nprint("Defined wait_for_job helper function")\n"
},
{
"type": "code",
- "source": "job_status = wait_for_job(\n poll_fn=lambda: client.customization.jobs.get_status(name=job.name),\n label=\"Training\",\n timeout_minutes=120,\n display_fn=training_progress,\n)",
+ "source": "job_status = wait_for_job(\n poll_fn=lambda: client.jobs.get_status(name=job.job.name, workspace=WORKSPACE),\n label=\"Training\",\n timeout_minutes=120,\n display_fn=training_progress,\n)\n\nif job_status.status != \"completed\":\n raise RuntimeError(f\"Training job finished with status: {job_status.status}\")",
"language": "python",
- "source_html": "job_status = wait_for_job(\n poll_fn=lambda: client.customization.jobs.get_status(name=job.name),\n label="Training",\n timeout_minutes=120,\n display_fn=training_progress,\n)\n"
+ "source_html": "job_status = wait_for_job(\n poll_fn=lambda: client.jobs.get_status(name=job.job.name, workspace=WORKSPACE),\n label="Training",\n timeout_minutes=120,\n display_fn=training_progress,\n)\n\nif job_status.status != "completed":\n raise RuntimeError(f"Training job finished with status: {job_status.status}")\n"
},
{
"type": "markdown",
- "source": "## 5. Verify Auto-Deployed Model\n\nSince we set `lora_enabled=True` in the customization job's `deployment_config`, the platform automatically creates a NIM deployment for the base model after training completes. The LoRA adapter is attached to the base model entity (enabled by default) and the deployment serves both the base weights and the adapter through a single NIM instance.",
- "source_html": "5. Verify Auto-Deployed Model
\nSince we set lora_enabled=True in the customization job's deployment_config, the platform automatically creates a NIM deployment for the base model after training completes. The LoRA adapter is attached to the base model entity (enabled by default) and the deployment serves both the base weights and the adapter through a single NIM instance.
\n"
+ "source": "## 5. Deploy Fine-Tuned Model\n\nAfter training completes, verify the LoRA adapter is attached to the base model entity, then create a NIM deployment with `lora_enabled=True` so both the base weights and adapter are served through a single deployment.",
+ "source_html": "5. Deploy Fine-Tuned Model
\nAfter training completes, verify the LoRA adapter is attached to the base model entity, then create a NIM deployment with lora_enabled=True so both the base weights and adapter are served through a single deployment.
\n"
},
{
"type": "code",
- "source": "ADAPTER_NAME = job.spec.output.name\nprint(f\"Looking for adapter: {ADAPTER_NAME}\")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f\"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)\")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f\"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s\"\n )\n\nprint(f\"Base model: {base_model.name}\")\nprint(f\"Adapter:\\n{adapter.model_dump_json(indent=2)}\")",
+ "source": "ADAPTER_NAME = OUTPUT_NAME\nprint(f\"Looking for adapter: {ADAPTER_NAME}\")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME, workspace=WORKSPACE)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f\"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)\")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f\"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s\"\n )\n\nprint(f\"Base model: {base_model.name}\")\nprint(f\"Adapter:\\n{adapter.model_dump_json(indent=2)}\")\n\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f\"tool-calling-deploy-cfg-{deploy_suffix}\"\ndeployment_name = f\"tool-calling-deploy-{deploy_suffix}\"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=WORKSPACE,\n name=DEPLOYMENT_CONFIG_NAME,\n engine=\"vllm\",\n model_spec={\n \"model_namespace\": WORKSPACE,\n \"model_name\": MODEL_NAME,\n \"lora_enabled\": True,\n },\n executor_config={\n \"gpu\": 1,\n \"image_name\": \"vllm/vllm-openai\",\n \"image_tag\": \"v0.22.1\",\n \"additional_args\": [\"--max-lora-rank\", \"32\"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace=WORKSPACE,\n name=deployment_name,\n config=deployment_config.name,\n)\n\nprint(f\"Deployment status: {deployment.status}\")",
"language": "python",
- "source_html": "ADAPTER_NAME = job.spec.output.name\nprint(f"Looking for adapter: {ADAPTER_NAME}")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s"\n )\n\nprint(f"Base model: {base_model.name}")\nprint(f"Adapter:\\n{adapter.model_dump_json(indent=2)}")\n"
+ "source_html": "ADAPTER_NAME = OUTPUT_NAME\nprint(f"Looking for adapter: {ADAPTER_NAME}")\n\n# The adapter may not be attached to the model entity immediately after\n# training completes — poll until it appears.\nADAPTER_TIMEOUT = 120\nadapter_start = time.time()\nadapter = None\nwhile time.time() - adapter_start < ADAPTER_TIMEOUT:\n base_model = client.models.retrieve(name=MODEL_NAME, workspace=WORKSPACE)\n matches = [a for a in (base_model.adapters or []) if a.name == ADAPTER_NAME]\n if matches:\n adapter = matches[0]\n break\n print(f"Adapter not yet attached, retrying... ({int(time.time() - adapter_start)}s)")\n time.sleep(5)\n\nif adapter is None:\n raise TimeoutError(\n f"Adapter '{ADAPTER_NAME}' not found on model '{MODEL_NAME}' within {ADAPTER_TIMEOUT}s"\n )\n\nprint(f"Base model: {base_model.name}")\nprint(f"Adapter:\\n{adapter.model_dump_json(indent=2)}")\n\ndeploy_suffix = uuid.uuid4().hex[:4]\nDEPLOYMENT_CONFIG_NAME = f"tool-calling-deploy-cfg-{deploy_suffix}"\ndeployment_name = f"tool-calling-deploy-{deploy_suffix}"\n\ndeployment_config = client.inference.deployment_configs.create(\n workspace=WORKSPACE,\n name=DEPLOYMENT_CONFIG_NAME,\n engine="vllm",\n model_spec={\n "model_namespace": WORKSPACE,\n "model_name": MODEL_NAME,\n "lora_enabled": True,\n },\n executor_config={\n "gpu": 1,\n "image_name": "vllm/vllm-openai",\n "image_tag": "v0.22.1",\n "additional_args": ["--max-lora-rank", "32"],\n },\n)\n\ndeployment = client.inference.deployments.create(\n workspace=WORKSPACE,\n name=deployment_name,\n config=deployment_config.name,\n)\n\nprint(f"Deployment status: {deployment.status}")\n"
},
{
"type": "markdown",
@@ -224,9 +226,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "DEPLOYMENT_NAME = f\"sft-deploy-{MODEL_NAME}\"\n\nTIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f\"Monitoring deployment '{DEPLOYMENT_NAME}'...\")\nprint(f\"Timeout: {TIMEOUT_MINUTES} minutes\\n\")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=DEPLOYMENT_NAME,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f\"Deployment: {DEPLOYMENT_NAME}\")\n print(f\"Status: {deployment_status.status}\")\n print(f\"Elapsed time: {elapsed_min}m {elapsed_sec}s\")\n\n if deployment_status.status == \"READY\":\n print(\"\\nDeployment is ready!\")\n if not client.models.wait_for_gateway(DEPLOYMENT_NAME, workspace=WORKSPACE, timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n\n if deployment_status.status in (\"FAILED\", \"ERROR\", \"TERMINATED\", \"LOST\", \"DELETED\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f\"Deployment timeout after {TIMEOUT_MINUTES} minutes\")\n\n time.sleep(15)",
+ "source": "TIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f\"Monitoring deployment '{deployment_name}'...\")\nprint(f\"Timeout: {TIMEOUT_MINUTES} minutes\\n\")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=deployment_name,\n workspace=WORKSPACE,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f\"Deployment: {deployment_name}\")\n print(f\"Status: {deployment_status.status}\")\n print(f\"Elapsed time: {elapsed_min}m {elapsed_sec}s\")\n\n if deployment_status.status in (\"RUNNING\", \"READY\"):\n print(\"\\nDeployment is ready!\")\n if not client.models.wait_for_gateway(deployment_name, workspace=WORKSPACE, timeout=60):\n raise RuntimeError(\"Inference gateway did not become ready\")\n break\n\n if deployment_status.status in (\"FAILED\", \"ERROR\", \"TERMINATED\", \"LOST\", \"DELETED\"):\n raise RuntimeError(f\"Deployment failed with status: {deployment_status.status}\")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f\"Deployment timeout after {TIMEOUT_MINUTES} minutes\")\n\n time.sleep(15)",
"language": "python",
- "source_html": "DEPLOYMENT_NAME = f"sft-deploy-{MODEL_NAME}"\n\nTIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f"Monitoring deployment '{DEPLOYMENT_NAME}'...")\nprint(f"Timeout: {TIMEOUT_MINUTES} minutes\\n")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=DEPLOYMENT_NAME,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f"Deployment: {DEPLOYMENT_NAME}")\n print(f"Status: {deployment_status.status}")\n print(f"Elapsed time: {elapsed_min}m {elapsed_sec}s")\n\n if deployment_status.status == "READY":\n print("\\nDeployment is ready!")\n if not client.models.wait_for_gateway(DEPLOYMENT_NAME, workspace=WORKSPACE, timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n\n if deployment_status.status in ("FAILED", "ERROR", "TERMINATED", "LOST", "DELETED"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f"Deployment timeout after {TIMEOUT_MINUTES} minutes")\n\n time.sleep(15)\n"
+ "source_html": "TIMEOUT_MINUTES = 30\nstart_time = time.time()\ntimeout_seconds = TIMEOUT_MINUTES * 60\n\nprint(f"Monitoring deployment '{deployment_name}'...")\nprint(f"Timeout: {TIMEOUT_MINUTES} minutes\\n")\n\nwhile True:\n deployment_status = client.inference.deployments.retrieve(\n name=deployment_name,\n workspace=WORKSPACE,\n )\n\n elapsed = time.time() - start_time\n elapsed_min = int(elapsed // 60)\n elapsed_sec = int(elapsed % 60)\n\n clear_output(wait=True)\n print(f"Deployment: {deployment_name}")\n print(f"Status: {deployment_status.status}")\n print(f"Elapsed time: {elapsed_min}m {elapsed_sec}s")\n\n if deployment_status.status in ("RUNNING", "READY"):\n print("\\nDeployment is ready!")\n if not client.models.wait_for_gateway(deployment_name, workspace=WORKSPACE, timeout=60):\n raise RuntimeError("Inference gateway did not become ready")\n break\n\n if deployment_status.status in ("FAILED", "ERROR", "TERMINATED", "LOST", "DELETED"):\n raise RuntimeError(f"Deployment failed with status: {deployment_status.status}")\n\n if elapsed > timeout_seconds:\n raise TimeoutError(f"Deployment timeout after {TIMEOUT_MINUTES} minutes")\n\n time.sleep(15)\n"
},
{
"type": "markdown",
@@ -235,9 +237,9 @@ export default { cells: [
},
{
"type": "code",
- "source": "test_messages = [\n {\"role\": \"user\", \"content\": \"Calculate the factorial of 12 using math functions.\"},\n]\n\ntest_tools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"math_factorial\",\n \"description\": \"Calculate the factorial of a given number.\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"number\": {\n \"type\": \"integer\",\n \"description\": \"The number for which factorial needs to be calculated.\",\n }\n },\n \"required\": [\"number\"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n \"\"\"Send a tool calling request and display the response.\"\"\"\n response = client.inference.gateway.model.post(\n \"v1/chat/completions\",\n name=model_name,\n body={\n \"messages\": test_messages,\n \"tools\": test_tools,\n \"tool_choice\": \"auto\",\n \"temperature\": 0,\n \"max_tokens\": 256,\n },\n )\n\n print(f\"{'=' * 60}\")\n print(f\" {label}\")\n print(f\"{'=' * 60}\")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, \"BASE MODEL (before fine-tuning)\")\ntest_tool_calling(ADAPTER_NAME, \"FINE-TUNED MODEL (after LoRA)\")",
+ "source": "test_messages = [\n {\"role\": \"user\", \"content\": \"Calculate the factorial of 12 using math functions.\"},\n]\n\ntest_tools = [\n {\n \"type\": \"function\",\n \"function\": {\n \"name\": \"math_factorial\",\n \"description\": \"Calculate the factorial of a given number.\",\n \"parameters\": {\n \"type\": \"object\",\n \"properties\": {\n \"number\": {\n \"type\": \"integer\",\n \"description\": \"The number for which factorial needs to be calculated.\",\n }\n },\n \"required\": [\"number\"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n \"\"\"Send a tool calling request and display the response.\"\"\"\n response = client.inference.gateway.provider.post(\n \"v1/chat/completions\",\n name=deployment_name,\n workspace=WORKSPACE,\n body={\n \"model\": model_name,\n \"messages\": test_messages,\n \"tools\": test_tools,\n \"tool_choice\": \"auto\",\n \"temperature\": 0,\n \"max_tokens\": 256,\n },\n )\n\n print(f\"{'=' * 60}\")\n print(f\" {label}\")\n print(f\"{'=' * 60}\")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, \"BASE MODEL (before fine-tuning)\")\ntest_tool_calling(ADAPTER_NAME, \"FINE-TUNED MODEL (after LoRA)\")",
"language": "python",
- "source_html": "test_messages = [\n {"role": "user", "content": "Calculate the factorial of 12 using math functions."},\n]\n\ntest_tools = [\n {\n "type": "function",\n "function": {\n "name": "math_factorial",\n "description": "Calculate the factorial of a given number.",\n "parameters": {\n "type": "object",\n "properties": {\n "number": {\n "type": "integer",\n "description": "The number for which factorial needs to be calculated.",\n }\n },\n "required": ["number"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n """Send a tool calling request and display the response."""\n response = client.inference.gateway.model.post(\n "v1/chat/completions",\n name=model_name,\n body={\n "messages": test_messages,\n "tools": test_tools,\n "tool_choice": "auto",\n "temperature": 0,\n "max_tokens": 256,\n },\n )\n\n print(f"{'=' * 60}")\n print(f" {label}")\n print(f"{'=' * 60}")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, "BASE MODEL (before fine-tuning)")\ntest_tool_calling(ADAPTER_NAME, "FINE-TUNED MODEL (after LoRA)")\n"
+ "source_html": "test_messages = [\n {"role": "user", "content": "Calculate the factorial of 12 using math functions."},\n]\n\ntest_tools = [\n {\n "type": "function",\n "function": {\n "name": "math_factorial",\n "description": "Calculate the factorial of a given number.",\n "parameters": {\n "type": "object",\n "properties": {\n "number": {\n "type": "integer",\n "description": "The number for which factorial needs to be calculated.",\n }\n },\n "required": ["number"],\n },\n },\n }\n]\n\n\ndef test_tool_calling(model_name: str, label: str):\n """Send a tool calling request and display the response."""\n response = client.inference.gateway.provider.post(\n "v1/chat/completions",\n name=deployment_name,\n workspace=WORKSPACE,\n body={\n "model": model_name,\n "messages": test_messages,\n "tools": test_tools,\n "tool_choice": "auto",\n "temperature": 0,\n "max_tokens": 256,\n },\n )\n\n print(f"{'=' * 60}")\n print(f" {label}")\n print(f"{'=' * 60}")\n print(json.dumps(response, indent=2))\n\n\ntest_tool_calling(MODEL_NAME, "BASE MODEL (before fine-tuning)")\ntest_tool_calling(ADAPTER_NAME, "FINE-TUNED MODEL (after LoRA)")\n"
},
{
"type": "markdown",
diff --git a/docs/fern/gated-nav.yml b/docs/fern/gated-nav.yml
index e210ce8bdc..b9d6128dd9 100644
--- a/docs/fern/gated-nav.yml
+++ b/docs/fern/gated-nav.yml
@@ -145,8 +145,3 @@
contents:
- page: Cluster Setup
path: ../../troubleshooting/cluster-setup.mdx
-- section: Fine-tune Models
- contents:
- # DPO is not yet a documented training type; the tutorial stays in the repo but out of nav.
- - page: DPO Customization Job
- path: ../../customizer/tutorials/dpo-customization-job.mdx
diff --git a/docs/fern/scripts/README.md b/docs/fern/scripts/README.md
index 75abc8bdde..ef43050b23 100644
--- a/docs/fern/scripts/README.md
+++ b/docs/fern/scripts/README.md
@@ -1,25 +1,21 @@
# Fern scripts
+Run these from the **repository root**. For local development, use `uv run` — dependencies are resolved via the workspace `pyproject.toml`.
+
+`requirements.txt` in this directory lists the direct Python dependencies for CI or other `pip install -r` workflows.
+
## `ipynb-to-fern-json.py`
Converts Jupyter notebooks to the JSON/TS format consumed by
`fern/components/NotebookViewer.tsx`. Pulled from
[NVIDIA-NeMo/DataDesigner](https://github.com/NVIDIA-NeMo/DataDesigner/blob/main/fern/scripts/ipynb-to-fern-json.py).
-### Setup
-
-```bash
-python3 -m venv .venv
-source .venv/bin/activate
-pip install -r fern/scripts/requirements.txt
-```
-
### Run
```bash
-python fern/scripts/ipynb-to-fern-json.py \
+uv run python docs/fern/scripts/ipynb-to-fern-json.py \
docs/customizer/tutorials/sft-customization-job.ipynb \
- -o fern/components/notebooks/sft-customization-job.json
+ -o docs/fern/components/notebooks/sft-customization-job.json
```
Writes both `.json` (canonical data) and `.ts` (default-export wrapper
@@ -37,3 +33,26 @@ After writing the `.ts` module, register it in `fern/components/NotebookViewer.t
colabUrl="https://colab.research.google.com/github/NVIDIA-NeMo/nemo-platform/blob/main/docs/customizer/tutorials/sft-customization-job.ipynb"
/>
```
+
+## `ipynb-to-mdx.py`
+
+Converts Jupyter notebooks to **inline Fern MDX** using `nemo_nb` (`NotebookConverter`,
+same engine as `nemo-nb to-sphinx-md`). Post-processes for Fern frontmatter, a Google
+Colab banner, and canonical `/documentation/...` internal links.
+
+### Run MDX conversion
+
+```bash
+uv run python docs/fern/scripts/ipynb-to-mdx.py --all-customizer-tutorials
+```
+
+Or a single notebook:
+
+```bash
+uv run python docs/fern/scripts/ipynb-to-mdx.py \
+ docs/customizer/tutorials/sft-customization-job.ipynb \
+ -o docs/customizer/tutorials/sft-customization-job.mdx \
+ --title "Full SFT Customization"
+```
+
+Re-run whenever the source `.ipynb` changes.
diff --git a/docs/fern/scripts/ipynb-to-fern-json.py b/docs/fern/scripts/ipynb-to-fern-json.py
index 9ca81d041f..5ce1128372 100644
--- a/docs/fern/scripts/ipynb-to-fern-json.py
+++ b/docs/fern/scripts/ipynb-to-fern-json.py
@@ -32,8 +32,19 @@
import json
import re
import sys
+from datetime import datetime
from pathlib import Path
+_CURRENT_YEAR = datetime.now().year
+_TS_FILE_HEADER = (
+ "/**\n"
+ f" * SPDX-FileCopyrightText: Copyright (c) 2025-{_CURRENT_YEAR} NVIDIA CORPORATION & AFFILIATES. All rights reserved.\n"
+ " * SPDX-License-Identifier: Apache-2.0\n"
+ " *\n"
+ " * Auto-generated by ipynb-to-fern-json.py - do not edit manually.\n"
+ " */\n"
+)
+
from markdown_it import MarkdownIt
from PIL import Image
from pygments import highlight
@@ -302,7 +313,7 @@ def write_ts_export(data: dict, ts_path: Path) -> None:
"""Write a .ts file that exports the notebook data inline (MDX imports the .ts, not the .json)."""
cells_json = json.dumps(data["cells"], indent=2, ensure_ascii=False)
ts_path.write_text(
- f"/** Auto-generated by ipynb-to-fern-json.py - do not edit */\nexport default {{ cells: {cells_json} }};\n",
+ f"{_TS_FILE_HEADER}export default {{ cells: {cells_json} }};\n",
encoding="utf-8",
)
diff --git a/docs/fern/scripts/ipynb-to-mdx.py b/docs/fern/scripts/ipynb-to-mdx.py
new file mode 100644
index 0000000000..f8841e6209
--- /dev/null
+++ b/docs/fern/scripts/ipynb-to-mdx.py
@@ -0,0 +1,178 @@
+#!/usr/bin/env python3
+# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
+# SPDX-License-Identifier: Apache-2.0
+
+"""Convert Jupyter notebooks to inline Fern MDX using nemo-nb.
+
+Wraps ``nemo_nb.converter.NotebookConverter`` (``nemo-nb to-sphinx-md``) and
+post-processes the output for Fern:
+ - Fern frontmatter (title, description)
+ - Google Colab link instead of the nemo-nb download anchor
+ - Relative doc links rewritten to canonical ``/documentation/...`` URLs
+
+Usage:
+ python ipynb-to-mdx.py input.ipynb -o output.mdx --title "Page Title"
+ python ipynb-to-mdx.py --all-customizer-tutorials
+"""
+
+from __future__ import annotations
+
+import re
+import sys
+from pathlib import Path
+
+from nemo_nb.converter import NotebookConverter
+
+COLAB_REPO = "https://colab.research.google.com/github/NVIDIA-NeMo/nemo-platform/blob/main"
+
+DOWNLOAD_LINK_RE = re.compile(
+ r'Download this tutorial as a Jupyter notebook\s*',
+ re.IGNORECASE,
+)
+
+_LINK_REWRITES: list[tuple[re.Pattern[str], str]] = [
+ (
+ re.compile(r"\]\(\.\./\.\./get-started/quickstart\.md?\)"),
+ "](/documentation/get-started)",
+ ),
+ (
+ re.compile(r"\]\(\.\./\.\./get-started/concepts/manage-secrets\.md?\)"),
+ "](/documentation/get-started/core-concepts/manage-secrets)",
+ ),
+ (
+ re.compile(r"\]\(\.\./manage-customization-jobs/hyperparameters\.md?\)"),
+ "](/documentation/customizer-reference/manage-customization-jobs/training-configuration)",
+ ),
+ (
+ re.compile(r"\]\(\.\./manage-customization-jobs/get-job-status\.md?\)"),
+ "](/documentation/customizer-reference/manage-customization-jobs/get-job-status)",
+ ),
+ (
+ re.compile(r"\]\(\.\./\.\./evaluator/index(?:\.md)?\)"),
+ "](/documentation/evaluate-models)",
+ ),
+ (
+ re.compile(r"\]\(\./distillation-customization-job(?:\.ipynb)?\)"),
+ "](/documentation/customizer-reference/tutorials/distillation-customization-job)",
+ ),
+ (
+ re.compile(r"\]\(\./embedding-customization-job(?:\.ipynb)?\)"),
+ "](/documentation/customizer-reference/tutorials/embedding-customization-job)",
+ ),
+ (
+ re.compile(r"\]\(\./lora-customization-job(?:\.ipynb)?\)"),
+ "](/documentation/customizer-reference/tutorials/lora-customization-job)",
+ ),
+ (
+ re.compile(r"\]\(\./optimize-throughput(?:\.ipynb)?\)"),
+ "](/documentation/customizer-reference/tutorials/optimize-throughput)",
+ ),
+ (
+ re.compile(r"\]\(\./sft-customization-job(?:\.ipynb)?\)"),
+ "](/documentation/customizer-reference/tutorials/sft-customization-job)",
+ ),
+ (
+ re.compile(r"\]\(fine-tune-metrics\)"),
+ "](/documentation/customizer-reference/tutorials/metrics)",
+ ),
+ (
+ re.compile(r"\]\(nemo-ms-about-concepts-customization\)"),
+ "](/documentation/customizer-reference/customization-concepts#nemo-ms-about-concepts-customization)",
+ ),
+]
+
+CUSTOMIZER_TUTORIALS: list[tuple[str, str]] = [
+ ("sft-customization-job.ipynb", "Full SFT Customization"),
+ ("lora-customization-job.ipynb", "LoRA Model Customization"),
+ ("distillation-customization-job.ipynb", "Knowledge Distillation Customization"),
+ ("embedding-customization-job.ipynb", "Embedding Model Customization"),
+ ("optimize-throughput.ipynb", "Optimize for Tokens/GPU Throughput"),
+]
+
+
+def rewrite_links(text: str) -> str:
+ for pattern, replacement in _LINK_REWRITES:
+ text = pattern.sub(replacement, text)
+ return text
+
+
+def repo_relative_path(path: Path) -> str:
+ repo_root = Path(__file__).resolve().parents[3]
+ return path.resolve().relative_to(repo_root).as_posix()
+
+
+def colab_link_for(ipynb_path: Path) -> str:
+ return f"[Run in Google Colab]({COLAB_REPO}/{repo_relative_path(ipynb_path)})"
+
+
+def convert_notebook_to_mdx(ipynb_path: Path, *, title: str) -> str:
+ body = NotebookConverter().convert(ipynb_path)
+ body = DOWNLOAD_LINK_RE.sub("", body).lstrip("\n")
+ body = rewrite_links(body)
+
+ return (
+ "---\n"
+ f'title: "{title}"\n'
+ 'description: ""\n'
+ "---\n"
+ "\n"
+ f"{colab_link_for(ipynb_path)}\n"
+ "\n"
+ f"{body.rstrip()}\n"
+ )
+
+
+def convert_all_customizer_tutorials(tutorials_dir: Path) -> int:
+ rc = 0
+ for notebook_name, title in CUSTOMIZER_TUTORIALS:
+ ipynb = tutorials_dir / notebook_name
+ mdx = tutorials_dir / notebook_name.replace(".ipynb", ".mdx")
+ if not ipynb.exists():
+ print(f"Error: {ipynb} not found", file=sys.stderr)
+ rc = 1
+ continue
+ mdx.write_text(convert_notebook_to_mdx(ipynb, title=title), encoding="utf-8")
+ print(f"Wrote {mdx}")
+ return rc
+
+
+def main() -> int:
+ args = sys.argv[1:]
+ if not args or "-h" in args or "--help" in args:
+ print(__doc__)
+ return 0
+
+ if "--all-customizer-tutorials" in args:
+ repo_root = Path(__file__).resolve().parents[3]
+ tutorials_dir = repo_root / "docs" / "customizer" / "tutorials"
+ return convert_all_customizer_tutorials(tutorials_dir)
+
+ input_path = Path(args[0])
+ output_path: Path | None = None
+ title: str | None = None
+
+ if "-o" in args:
+ idx = args.index("-o")
+ if idx + 1 < len(args):
+ output_path = Path(args[idx + 1])
+ if "--title" in args:
+ idx = args.index("--title")
+ if idx + 1 < len(args):
+ title = args[idx + 1]
+
+ if not input_path.exists():
+ print(f"Error: {input_path} not found", file=sys.stderr)
+ return 1
+ if output_path is None:
+ output_path = input_path.with_suffix(".mdx")
+ if not title:
+ print("Error: --title is required", file=sys.stderr)
+ return 1
+
+ output_path.write_text(convert_notebook_to_mdx(input_path, title=title), encoding="utf-8")
+ print(f"Wrote {output_path}")
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/docs/get-started/concepts/filtering.mdx b/docs/get-started/concepts/filtering.mdx
index 2ac7bd2f27..304b428b34 100644
--- a/docs/get-started/concepts/filtering.mdx
+++ b/docs/get-started/concepts/filtering.mdx
@@ -180,5 +180,8 @@ metrics = client.evaluation.metrics.list(
### Filtering by status
```python
-jobs = client.customization.jobs.list(filter={"status": "completed"})
+jobs = client.jobs.list(
+ workspace="default",
+ filter={"source": "automodel", "status": "completed"},
+)
```
diff --git a/docs/troubleshooting/customizer.mdx b/docs/troubleshooting/customizer.mdx
index 3b90b8b408..c5a994e343 100644
--- a/docs/troubleshooting/customizer.mdx
+++ b/docs/troubleshooting/customizer.mdx
@@ -5,7 +5,7 @@ description: ""
**Job fails during model download:**
- Verify the HuggingFace token secret is configured correctly
- Accept the model's license on the [HuggingFace model page](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)
-- Check job status: `client.customization.jobs.retrieve(name=job.name, workspace="default")`
+- Check job status: `client.jobs.get_status(name=job.name, workspace="default")`
**Job fails with disk full or 500 error when retrieving logs:**
- The platform's shared persistent volume is likely full. Customization jobs require significant disk space: ~3× model size for full SFT, ~1.5× for LoRA. If you are also deploying the model from a base checkpoint fileset, plan for ~2.5× model size overall.
diff --git a/packages/nmp_testing/src/nmp/testing/e2e/customizer.py b/packages/nmp_testing/src/nmp/testing/e2e/customizer.py
index c0a9180b5e..ff2f837d59 100644
--- a/packages/nmp_testing/src/nmp/testing/e2e/customizer.py
+++ b/packages/nmp_testing/src/nmp/testing/e2e/customizer.py
@@ -218,7 +218,7 @@ def get_job_failure_details(sdk: NeMoPlatform, job_name: str, workspace: str) ->
def _log_training_progress(status) -> None:
"""Extract and log training progress from the job status steps structure."""
for job_step in status.steps or []:
- if job_step.name == "customization-training-job":
+ if job_step.name == "training":
for task in job_step.tasks or []:
task_details = task.status_details or {}
step = task_details.get("step")
diff --git a/plugins/nemo-automodel/src/nemo_automodel_plugin/transform.py b/plugins/nemo-automodel/src/nemo_automodel_plugin/transform.py
index d9bf668f05..a3a3c70647 100644
--- a/plugins/nemo-automodel/src/nemo_automodel_plugin/transform.py
+++ b/plugins/nemo-automodel/src/nemo_automodel_plugin/transform.py
@@ -41,11 +41,6 @@ async def transform_input_to_output(
await check_dataset_access(sdk, input_spec.dataset.validation, workspace)
is_embedding = bool(model_entity.spec and getattr(model_entity.spec, "is_embedding_model", False))
- if is_embedding:
- raise ValueError(
- "Embedding-model SFT is not supported in Automodel v1. "
- "Use a causal LM checkpoint or wait for a future release."
- )
output_type = _infer_output_type(input_spec, is_embedding)
diff --git a/services/automodel/src/nmp/automodel/tasks/training/backends/checkpoints.py b/services/automodel/src/nmp/automodel/tasks/training/backends/checkpoints.py
index f43220fe2f..770fd73aae 100644
--- a/services/automodel/src/nmp/automodel/tasks/training/backends/checkpoints.py
+++ b/services/automodel/src/nmp/automodel/tasks/training/backends/checkpoints.py
@@ -365,7 +365,7 @@ def export_onnx(
tokenizer_path=tokenizer_path,
pooling="avg",
normalize=True,
- opset=17,
+ opset=18,
export_dtype="fp16",
verify=True,
)
diff --git a/services/core/models/src/nmp/core/models/api/v2/models.py b/services/core/models/src/nmp/core/models/api/v2/models.py
index aa84d33ffd..4a8d80fdef 100644
--- a/services/core/models/src/nmp/core/models/api/v2/models.py
+++ b/services/core/models/src/nmp/core/models/api/v2/models.py
@@ -269,6 +269,10 @@ async def start_update_model_spec_job(model_entity: ModelEntity):
name="model-spec-analysis",
executor=CPUExecutionProviderSpec(
provider="cpu",
+ # Profile ``gpu`` selects the Docker CPU executor (see jobs.executors
+ # in platform config). Avoid ``default``, which is translated to the
+ # host subprocess backend when subprocess/default is registered.
+ profile="gpu",
container=ContainerSpec(
image=get_qualified_image("nmp-automodel-tasks"),
entrypoint=["/opt/venv/bin/python"],