Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 0 additions & 10 deletions 3.test_cases/8.neuronx-nemo-megatron/1.convert-weight.sbatch

This file was deleted.

30 changes: 30 additions & 0 deletions 3.test_cases/8.neuronx-nemo-megatron/1.setup-venv.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
#!/usr/bin/env bash
set -euxo pipefail
APPS_PATH="$1"

cd ${APPS_PATH}
# Install Python venv
sudo apt-get install -y python3.8-venv g++

# Create Python venv
python3.8 -m venv aws_neuron_venv_pytorch

# Activate Python venv
source ${APPS_PATH}/aws_neuron_venv_pytorch/bin/activate
python -m pip install -U pip

# Install Jupyter notebook kernel
pip install ipykernel
python3.8 -m ipykernel install --user --name aws_neuron_venv_pytorch --display-name "Python (torch-neuronx)"
pip install jupyter notebook
pip install environment_kernels

# Set pip repository pointing to the Neuron repository
python -m pip config set global.extra-index-url https://pip.repos.neuron.amazonaws.com

# Install wget, awscli
python -m pip install wget
python -m pip install awscli

# Install Neuron Compiler and Framework
python -m pip install neuronx-cc==2.* torch-neuronx torchvision
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
#!/usr/bin/env bash
set -euxo pipefail
APPS_PATH="$1"

cd ${APPS_PATH}
source ${APPS_PATH}/aws_neuron_venv_pytorch/bin/activate
git clone https://github.com/aws-neuron/neuronx-nemo-megatron.git
cd neuronx-nemo-megatron
pip3 install wheel
./build.sh
pip3 install ./build/*.whl
pip3 install -r requirements.txt torch==1.13.1 protobuf==3.20.3

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Versions could be provided as arguments

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actually, the neuronx-nemo-megatron is picky for the pytorch version. I would like to hard code the value until it relax the requirement.

# You also need following to run the weight conversion
pip install accelerate
python3 -c "from nemo.collections.nlp.data.language_modeling.megatron.dataset_utils import compile_helper; \
compile_helper()"
17 changes: 0 additions & 17 deletions 3.test_cases/8.neuronx-nemo-megatron/2.tokenize.sh

This file was deleted.

13 changes: 13 additions & 0 deletions 3.test_cases/8.neuronx-nemo-megatron/3.convert-weight.sbatch
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
#!/bin/bash
#SBATCH --exclusive
#SBATCH --output=slurm-%x-%j.out
#SBATCH --cpus-per-task 96
#SBATCH --nodes 1

: "${APPS_PATH:=/fsx}"
: "${MODEL_PATH:=/fsx}"
: "${DATA_PATH:=/fsx/data/books}"

source ${APPS_PATH}/aws_neuron_venv_pytorch/bin/activate
python ${APPS_PATH}/aws_neuron_venv_pytorch/lib/python3.8/site-packages/transformers/models/llama/convert_llama_weights_to_hf.py \
--input_dir ${MODEL_PATH}/Llama2-meta --model_size 7B --output_dir ${MODEL_PATH}/Llama2-7b-hf
5 changes: 0 additions & 5 deletions 3.test_cases/8.neuronx-nemo-megatron/3.precompile-model.sh

This file was deleted.

5 changes: 0 additions & 5 deletions 3.test_cases/8.neuronx-nemo-megatron/4.pretrain-model.sh

This file was deleted.

20 changes: 20 additions & 0 deletions 3.test_cases/8.neuronx-nemo-megatron/4.tokenize.sbatch
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
#!/bin/bash
#SBATCH --exclusive
#SBATCH --output=slurm-%x-%j.out
#SBATCH --cpus-per-task 128
#SBATCH --nodes 1

: "${APPS_PATH:=/fsx}"
: "${MODEL_PATH:=/fsx}"
: "${DATA_PATH:=/fsx/data/books}"
source ${APPS_PATH}/aws_neuron_venv_pytorch/bin/activate
python ${APPS_PATH}/neuronx-nemo-megatron/nemo/scripts/nlp_language_modeling/preprocess_data_for_megatron.py \
--input=${DATA_PATH}/book.jsonl \
--json-keys=text \
--tokenizer-library=huggingface \
--tokenizer-type=${MODEL_PATH}/Llama2-7b-hf \
--dataset-impl=mmap \
--output-prefix=${DATA_PATH}/book-tokenized \
--append-eod \
--need-pad-id \
--workers=128
4 changes: 4 additions & 0 deletions 3.test_cases/8.neuronx-nemo-megatron/5.precompile-model.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
#!/bin/bash
cd ${APPS_PATH}/neuronx-nemo-megatron/nemo/examples/nlp/language_modeling

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Line 2-6 is redundant to the readme file where variables are set. Remove these lines and integrate 7-9 directly in the read-me file (3 lines) as part of the instructions. That'll remove one file and make it simple to follow.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

True, omitted.

source ${APPS_PATH}/aws_neuron_venv_pytorch/bin/activate
sbatch --cpus-per-task 1 --nodes 4 --output ${TEST_CASE_PATH}/slurm-%x-%j.out compile.slurm ./llama_7b.sh
4 changes: 4 additions & 0 deletions 3.test_cases/8.neuronx-nemo-megatron/6.pretrain-model.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
#!/bin/bash
cd ${APPS_PATH}/neuronx-nemo-megatron/nemo/examples/nlp/language_modeling
source ${APPS_PATH}/aws_neuron_venv_pytorch/bin/activate
sbatch --nodes 4 --output ${TEST_CASE_PATH}/slurm-%x-%j.out run.slurm ./llama_7b.sh
120 changes: 88 additions & 32 deletions 3.test_cases/8.neuronx-nemo-megatron/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,13 +5,38 @@
## 1. Preparation

This guide assumes that you have the following:
* A functional Slurm cluster on AWS.
* Neuron SDK and Torch-neuronx installed.
* A functional Slurm cluster on AWS. We also assume that Ubuntu AMI is used.
* Neuron SDK is installed on the cluster (see [AWS Neuron SDK documentation](https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/setup/torch-neuronx.html#setup-torch-neuronx) for the steps).
* An FSx for Lustre filesystem mounted on `/fsx`.
* `torch-neuronx` environment set up as virtual environment as `aws_neuron_venv_pytorch`. See [NeuronSDK documentation](https://awsdocs-neuron.readthedocs-hosted.com/en/latest/general/setup/neuron-setup/pytorch/neuronx/ubuntu/torch-neuronx-ubuntu20.html#setup-torch-neuronx-ubuntu20) for the setup.
* `neuronx-nemo-megatron` cloned on home directory of the slurm headnode (`cd ~ && git clone https://github.com/aws-neuron/neuronx-nemo-megatron.git`).

We recommend that you setup a Slurm cluster using the template in the architectures directory.
We recommend that you setup a Slurm cluster using the template in the architectures directory. Before creating the Slurm cluster, you need to setup the following environment variables:

```bash
export APPS_PATH=/fsx
export FSX_PATH=/fsx
export MODEL_PATH=/fsx
export DATA_PATH=$FSX_PATH/data/books
export TEST_CASE_PATH=${APPS_PATH}/awsome-distributed-training/3.test_cases/8.neuronx-nemo-megatron # where you copy the test case or set to your test case path
```

1. First of all, you need to have a Python virtual environment for `torch-neuronx` under `APPS_PATH`.

```bash
bash 1.setup-venv.sh ${APPS_PATH} # The argument specifies APPS_PATH
```

2. `neuronx-nemo-megatron` library need to be installed (and initialized) in the environment.


```bash
bash 2.setup-neuronx-nemo-megatron.sh ${APPS_PATH} #
```
You will see the following ERROR line during the script execution. This is safe to ignore.

```console
+ python3 -c 'from nemo.collections.nlp.data.language_modeling.megatron.dataset_utils import compile_helper; compile_helper()'
2023-Nov-18 09:17:45.728072 175272:175272 ERROR TDRV:tdrv_get_dev_info No neuron device available
```

## 1. Prepare Llama2 model

Expand All @@ -21,7 +46,7 @@ You can submit access request from [here](https://ai.meta.com/resources/models-a
We will assume that you had placed the model and tokenizer as follows on cluster:

```
/fsx/Llama2-meta/
${MODEL_PATH}/Llama2-meta/
├── 7B/
│ ├── checklist.chk
│ ├── consolidated.00.pth
Expand All @@ -33,13 +58,26 @@ We will assume that you had placed the model and tokenizer as follows on cluster
To convert the model to the standard Hugging Face format, the following script in transformers can be called with the following (example) command:

```
sbatch 1.convert-weight.sbatch
sbatch 3.convert-weight.sbatch
```

Note: For the purposes of this sample we assume you have saved the Llama-2-7b model in a directory called `Llama2-7b-hf` with the following format:
You can check progress of with `tail` command.

```

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

console for markdown formatting

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for catching this! updated

$ tail -f slurm-3.convert-weight.sbatch-xxx.out
```

```console
Fetching all parameters from the checkpoint at /fsx/Llama2-meta/7B.
Loading the checkpoint in a Llama model.
Loading checkpoint shards: 100%|██████████| 33/33 [00:12<00:00, 2.65it/s]
...
```
/fsx/Llama2-7b-hf/

Once the job completed, you will have the Llama-2-7b model weights and tokenizer in a huggingface format under a directory called `Llama2-7b-hf` with the following format:

```console
${DATAPATH}/Llama2-7b-hf/
├── config.json
├── generation_config.json
├── pytorch_model-00001-of-00002.bin
Expand All @@ -54,16 +92,16 @@ Note: For the purposes of this sample we assume you have saved the Llama-2-7b mo
## 2. Download and Tokenize dataset
This tutorial makes use of a [Red pyjama dataset](https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T). The dataset can be downloaded to your cluster by running the following commands on the head node:

```
mkdir -p /fsx/data/llama2
wget https://data.together.xyz/redpajama-data-1T/v1.0.0/book/book.jsonl # Note: Dataset download is 50G and will take approximately 3-4 hours to download.
or
wget https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T-Sample/resolve/main/book_sample.jsonl -O /fsx/data/llama2/book.jsonl
```bash
mkdir -p ${DATA_PATH}
wget https://data.together.xyz/redpajama-data-1T/v1.0.0/book/book.jsonl -O ${DATA_PATH}/book.jsonl # Note: Dataset download is 50G and will take approximately 3-4 hours to download. You can also use https://aria2.github.io/ for faster download
# or
# wget https://huggingface.co/datasets/togethercomputer/RedPajama-Data-1T-Sample/resolve/main/book_sample.jsonl -O ${DATA_PATH}/book.jsonl # Smaller sample dataset for quick testing
```

Once you have the Tokenizer and the dataset. You can tokenize the dataset following the below command:
```
sbatch 2.tokenize.sbatch
```bash
sbatch 4.tokenize.sbatch
```

Post tokenizing the dataset, you will have a path to the tokenizer and the dataset which will be used for pretraining.
Expand Down Expand Up @@ -94,26 +132,33 @@ An alternative to the JIT flow is to use the included [neuron_parallel_compile](

Before starting the compilation you need to update your path to the dataset and tokenizer in the llama_7b script as below :

```bash
cd ${APPS_PATH}/neuronx-nemo-megatron/nemo/examples/nlp/language_modeling
vi test_llama.sh
```
cd ~/neuronx-nemo-megatron/nemo/examples/nlp/language_modeling
vi llama_7b.sh
```
Update the below lines to
```
# For tokenizer
model.tokenizer.type='/fsx/Llama2-7b-hf' \

# For Dataset
model.data.data_prefix=[1.0,/fsx/data/books/book.jsonl-processed_text_document] \
Update the below lines

```bash
: ${TOKENIZER_PATH=$HOME/llamav2_weights/7b-hf}
: ${DATASET_PATH=$HOME/examples_datasets/llama_7b/book.jsonl-processed_text_document}
```

Run the following command to launch an AOT pre-compilation job on your ParallelCluster:
to
```bash
: ${TOKENIZER_PATH=${MODEL_PATH}/Llama2-7b-hf}
: ${DATASET_PATH=${DATA_PATH}/book-tokenized_text_document}
```
bash 3.precompile-model.sh

Then, run the following command to launch an AOT pre-compilation job on your ParallelCluster:

```bash
bash 5.precompile-model.sh
```

Once you have launched the precompilation job, run the `squeue` command to view the SLURM job queue on your cluster. If you have not recently run a job on your cluster, it may take 4-5 minutes for the requested trn1.32xlarge nodes to be launched and initialized. Once the job is running, `squeue` should show output similar to the following:
```

```console
JOBID PARTITION NAME USER ST TIME NODES NODELIST(REASON)
10 compute1 compile.slurm ubuntu R 5:11 4 compute1-dy-queue1-i1-[1-4]
```
Expand All @@ -124,7 +169,8 @@ tail -f slurm-compile.slurm-10.out
```

Once the precompilation job is complete, you should see a message similar to the following in the logs:
```

```console
2023-06-11 23:04:08.000738: INFO ||PARALLEL_COMPILE||: Total graphs: 22
2023-06-11 23:04:08.000738: INFO ||PARALLEL_COMPILE||: Total successful compilations: 22
2023-06-11 23:04:08.000738: INFO ||PARALLEL_COMPILE||: Total failed compilations: 0
Expand All @@ -137,18 +183,28 @@ At this point, you can press `CTRL-C` to exit the tail command.
Submit the training job

```
bash 4.pretrain-model.sh
bash 6.pretrain-model.sh
```


As outlined above, you can again use the `squeue` command to view the job queue. Once you see that your pretraining job is running, you can view the output of the training job by examining the file named `slurm-run.slurm-ZZ.out` where ZZ represents the JOBID of your job:
```

```bash
tail -f slurm-run.slurm-11.out
```

Once the model is loaded onto the Trainium accelerators and training has commenced, you will begin to see output indicating the job progress:
```

```console
Epoch 0: 22%|██▏ | 4499/20101 [22:26:14<77:48:37, 17.95s/it, loss=2.43, v_num=5563, reduced_train_loss=2.470, gradient_norm=0.121, parameter_norm=1864.0, global_step=4512.0, consumed_samples=1.16e+6, iteration_time=16.40]
Epoch 0: 22%|██▏ | 4500/20101 [22:26:32<77:48:18, 17.95s/it, loss=2.43, v_num=5563, reduced_train_loss=2.470, gradient_norm=0.121, parameter_norm=1864.0, global_step=4512.0, consumed_samples=1.16e+6, iteration_time=16.40]
Epoch 0: 22%|██▏ | 4500/20101 [22:26:32<77:48:18, 17.95s/it, loss=2.44, v_num=5563, reduced_train_loss=2.450, gradient_norm=0.120, parameter_norm=1864.0, global_step=4512.0, consumed_samples=1.16e+6, iteration_time=16.50]
```

## 5. Authors / Reviewers

* [A] Keita Watanabe - mlkeita@
* [R] Verdi March - marcverd@
* [R] Brad Doran
* [R] Justin Pirtle
* [R] Pierre-Yves Aquilanti - pierreya@