Commits · 12ff0e239bb4ca81b1d416ec0476d71179f36b94 · OpenDAS / apex

09 Aug, 2022 7 commits
- Merge remote-tracking branch 'origin/dev/hubertlu/flaky_tests' into IFU-master-2022-07-29 · 12ff0e23
  hubertlu-tw authored Aug 09, 2022
  
  12ff0e23
- Remove some comments in run_test.py · cebbb04f
  hubertlu-tw authored Aug 09, 2022
  
  cebbb04f
- Merge remote-tracking branch 'origin/dev/hubertlu/flaky_tests' into IFU-master-2022-07-29 · f1f28ff6
  hubertlu-tw authored Aug 09, 2022
  
  f1f28ff6
- Remove run_pyprof_data and run_pyprof_nvtx unit tests · 4d567459
  hubertlu-tw authored Aug 09, 2022
  
  4d567459
- Update L0 unit test script · ced59fcc
  hubertlu-tw authored Aug 09, 2022
  
  ced59fcc
- Skip a flaky unit test · 8a8eb34f
  hubertlu-tw authored Aug 09, 2022
  
  8a8eb34f
- Skip some flaky unit tests · 975a0e53
  hubertlu-tw authored Aug 09, 2022
  
  975a0e53
08 Aug, 2022 6 commits
- Un-skip some tests and skip some flaky tests · 1b7b02ef
  hubertlu-tw authored Aug 08, 2022
  
  1b7b02ef
- Addd a wrapper to skip flaky unit tests. · 4cfbe05c
  hubertlu-tw authored Aug 08, 2022
  
  4cfbe05c
- Skip the failing unit tests from the FusedRMSNorm PR (#85) · 87fc4125
  Hubert Lu authored Aug 08, 2022
```
* Skip the failing unit tests from the FusedRMSNorm PR

* Update test_lamb.py
Co-authored-by: Jithun Nair <37884920+jithunnair-amd@users.noreply.github.com>
```
  87fc4125
- Fix the cuda-specific transformer utils for ROCm · 57dea7f2
  hubertlu-tw authored Aug 08, 2022
  
  57dea7f2
- Merge remote-tracking branch 'origin/master' into IFU-master-2022-07-29 · cb8b7a88
  hubertlu-tw authored Aug 08, 2022
  
  cb8b7a88
- Revert code changes to mutltihead_attn tests · 51783cc7
  hubertlu-tw authored Aug 08, 2022
  
  51783cc7
05 Aug, 2022 1 commit

Hubert Lu authored Aug 05, 2022



* FusedRMSNorm/"T5LayerNorm" based on FusedLayerNorm (#1274)

* FusedRMSNorm based on FusedLayerNorm

* refactor duplicated kernels

* delete comments

* delete comments

* cleanup

* cleanup

* cleanup, fixed clobbering forward_affine_mixed_dtypes

* fix pybind naming and add MixedFused test

* undo skipping

* check elementwise_affine

* Update tests/L0/run_fused_layer_norm/test_fused_layer_norm.py

Oof, nice catch, thanks
Co-authored-by: Masaki Kozuki <masaki.kozuki.2014@gmail.com>
Co-authored-by: Masaki Kozuki <masaki.kozuki.2014@gmail.com>

* fix and generate docs for FusedRMSNorm (#1285)

* [FusedRMSNorm doc] document where epsilon is added (#1295)

* [FusedRMSNorm doc] add epsilon to formula

* correct

* better wording

* Fix some bugs

* Optimize HostRMSNormGradient and HostApplyRMSNorm for AMD GPUs

* Fix NaN issues in FusedRMSNorm

* Update test_fused_layer_norm.py

* Skip test_fused_layer_norm.TestAutocastFusedRMSNorm on ROCm

* Use at::cuda::warp_size() instead of at::cuda::getCurrentDeviceProperties()->warpSize
Co-authored-by: eqy <eddiey@nvidia.com>
Co-authored-by: Masaki Kozuki <masaki.kozuki.2014@gmail.com>
Co-authored-by: Stas Bekman <stas00@users.noreply.github.com>

c97ebfab

29 Jul, 2022 5 commits
- Fix some compiling errors · 038ed999
  hubertlu-tw authored Jul 29, 2022
  
  038ed999
- Unskip run_transformer unit tests · bbf2c8d0
  hubertlu-tw authored Jul 29, 2022
  
  bbf2c8d0
- Merge remote-tracking branch 'upstream/master' into IFU-master-2022-07-29 · 795a5e5b
  hubertlu-tw authored Jul 29, 2022
  
  795a5e5b
- Merge remote-tracking branch 'origin/dev/hubertlu/FusedRMSNorm' · 016c8d4f
  hubertlu-tw authored Jul 29, 2022
  
  016c8d4f
- Update test_fused_layer_norm.py · 0df6c4c3
  hubertlu-tw authored Jul 29, 2022
  
  0df6c4c3
28 Jul, 2022 1 commit

Sequence parallel perf updates (#1437) · 3c19f106

Eric Harper authored Jul 28, 2022



* use _all_gather_base
Signed-off-by: ericharper <complex451@gmail.com>

* use _reduce_scatter_base
Signed-off-by: ericharper <complex451@gmail.com>

* remove torch empty in backward
Signed-off-by: ericharper <complex451@gmail.com>

* check self.attn_mask_type
Signed-off-by: ericharper <complex451@gmail.com>

* remove extra arg
Signed-off-by: ericharper <complex451@gmail.com>

* update get_tensor_shapes logic
Signed-off-by: ericharper <complex451@gmail.com>

3c19f106

26 Jul, 2022 2 commits

Improvements in distributed Adam optimizer for Megatron (#1432) · 2e025ab5

Tim Moon authored Jul 26, 2022

* Improvements in distributed Adam optimizer for Megatron

Add option to allocate gradient buckets out of one large buffer. Add option to initialize params in user-provided order. Perform communication when saving optimizer state. Support param sync with any dtype.

* Style fixes in distributed Adam helper classes

Review suggestions from @crcrpar

2e025ab5

Fix bug when initializing model-parallel process groups for GPT-3 (#1435) · fb21698e

Tim Moon authored Jul 26, 2022

* Hack to enable training GPT-3

Seems to fix bug from #1416

* Add test to initialize model-parallelism for decoder-only Transformers

Namely GPT-3.

fb21698e

25 Jul, 2022 1 commit
- [transformer] update tests (#1428) · e57d9e79
  Aidyn-A authored Jul 25, 2022
  
  e57d9e79
21 Jul, 2022 2 commits
- Merge pull request #1429 from NVIDIA/update_spatial_bottleneck · 208d9670
  Thor Johnsen authored Jul 21, 2022
```
Bug fixes, perf improvements
```
  208d9670
- Bug fixes, perf improvements · f687e7fa
  Thor Johnsen authored Jul 21, 2022
  
  f687e7fa
20 Jul, 2022 1 commit

[transformer] UCC async test (#1417) · a29a698f

Aidyn-A authored Jul 20, 2022

* add test

* update batch sizes

* update batch sizes

* small updates

* delete comment

* add async comm

* add sync if needed

* update tests

* remove redundant imports

* code cleanup

* minor updates

* update dtype for comparison

* fix dtypes

* fix typo

* modify sizes and use common_utils.find_free_port

* fix typo and use double precision

* revert some changes, create test for profiling on L1

* remove redundant line

* revert UCC_TLS and add sync to fwd_bwd

* code clean up

* code clean up

* modify BERT test

* add comment

a29a698f

14 Jul, 2022 2 commits

[contrib] Fix the reference implementation of multihead_attn (#1423) · 809043f5

Masaki Kozuki authored Jul 14, 2022



* follow the current signature
Signed-off-by: Masaki Kozuki <mkozuki@nvidia.com>

* call .backward on outputs
Signed-off-by: Masaki Kozuki <mkozuki@nvidia.com>

* update the other caller of _softmax_backward_data
Signed-off-by: Masaki Kozuki <mkozuki@nvidia.com>

809043f5

Time dimension shape check for fused scale mask softmax kernel (#1421) · 1337e81e

Sandeep Subramanian authored Jul 13, 2022



* Time dimension shape check for fused scale mask softmax kernel
Signed-off-by: MaximumEntropy <sandeep.subramanian.1@umontreal.ca>

* Add shape test
Signed-off-by: MaximumEntropy <sandeep.subramanian.1@umontreal.ca>

* Fix mask shape
Signed-off-by: MaximumEntropy <sandeep.subramanian.1@umontreal.ca>

1337e81e

11 Jul, 2022 1 commit

update: mpu for t5 rpe (#1416) · 5ff5a884

Perkz Zheng authored Jul 12, 2022



* update: mpu for t5 rpe

* update: add rpe mpu group test

* fix semicolon bugs
Co-authored-by: Masaki Kozuki <masaki.kozuki.2014@gmail.com>

* fix semicolon bugs
Co-authored-by: Masaki Kozuki <masaki.kozuki.2014@gmail.com>
Co-authored-by: Masaki Kozuki <masaki.kozuki.2014@gmail.com>

5ff5a884

07 Jul, 2022 1 commit
- Remove `pyprof` and `reparameterization` (#1404) · 8a7a3325
  Masaki Kozuki authored Jul 06, 2022
```
* remove pyprof

* remove reparameterization

* remove pyprof test

* clean up
```
  8a7a3325
05 Jul, 2022 2 commits

Add features to distributed Adam for Megatron support (#1414) · cd499737

Tim Moon authored Jul 05, 2022

* Add features to distributed Adam for Megatron support

Support gradient clipping, gradient scaling, FP32 grad accumulation, and multiple dtypes and devices.

* Restore closure arg to distributed Adam

Review suggestion from @crcrpar

cd499737

[UCC][TORCH_UCC]Do integer driver version comparison for UCC (#1411) · bf3c008e
eqy authored Jul 05, 2022
```
* Integer driver number comparison

* packaging
```
bf3c008e

23 Jun, 2022 2 commits

[transformer] Port Sequence Parallelism (takeover of #1396) (#1400) · 3ff1a10f

Masaki Kozuki authored Jun 23, 2022

* it looks possible to remove this file

* add communication collectives

* update Column|RowParallelLinear

* update checkpoint function

* update function name

* parity between public and private collectives

* row parallel linear

* column parallel linear

* sequence parallel: p2p comm

fix typo

* sequence parallel: pipeline parallel

* fix typo

* add layernorm with sequence_parallel_enabled attr

* class variable -> member variable

* fix col parallel test with sequence parallel

* Initial test of `forward_backward_pipelining_without_interleaving` with `model_type=ModelType.encoder_and_decoder`

* add cases pretending to test sequence_parallel

* Apply 2 suggestion(s) to 1 file(s)

* update sequence_parallel_enabled docstring

* update docstring: order of tensor dimensions, sequence_parallel_enabled behavior

* Divide sequence_length if sequence parallel

tensor shape should be updated if sequence parallel is enabled.

* cherry-pick https://github.com/NVIDIA/Megatron-LM/commit/8474e6e54fcb9dfa37aea039352f9fb485fb6f61

* type annotation

* Fix matmul call in RowParallelLinear

Fix `sequence_parallel_enabled` to `False` as you can see in
https://github.com/NVIDIA/Megatron-LM/blob/d898a8991d1a08d29074f87819d1bf41517e35f5/megatron/mpu/layers.py#L511-L514

* update rowparallellinear test

* fix `loss_weight` is not defined in test_layers

* @eqy's comment

* mixed fused layer norm

* fix typo

* misc

* test_layers cleanup

* Skip Bert/GPT script

Since these two models haven't gotten updated for sequence parallle, e.g. the update of the order of dimension from (batch, sequence, feature) to (sequence, batch, feature) and global variables of arguments

* debug part 1/N: comment out `x.retain_grad`

* debug part 2/N: [ColumnParallelLinear] comment out overriding of sequence_parallel_enabled

* debug 3/N: add pipeline test with parallel mlp

* Fix handling `self.input_tensor` and argument

* tp2pp4 ModelType.encoder_or_decoder is failing, which can be at my fault because the backward is blaming the output and the grad_ouptut shape don't match

* revert debug 1/N

* defer tensor model parallel size > 1

* split tensor in sequence dim

* cosmetic

* cosmetic: remove archaic comment

* enable TP>1 for encoder_and_decoder as well

* set requires_grad=True always...

* Set `scatter_gather_tensors_in_pipeline` to :obj:`False`

for the sake of nemo megatron's GPT works with sequence parallel enabled.

* brush up comment of `requires_grad()`

There's a possibility that PyTorch DistributedDataParallel hangs
when some tensor (or parameter) doesn't require grad according to @ptrblck.
This forced `requires_grad` in my understanding is different from that.

* misc changes of scatter_gather_tensors_in_pipeline comment

* guard for torch_ucc

* cosmetic changes related to tests

* update command line arguments

* update TransformerLanguageModel

* rename

* move gpt to gpt.py

* update bert

* add all_gather for params in sequence parallel region

* misc. some diffs were lost during rebasing...

* updates for non sequence parallel execution

* gpt with sequence parallel

* Apply 2 suggestion(s) to 2 file(s)

* update tensor&pipeline parallel size

* why `sequence_parallel_enabled` is not supplied!? Did I messed up when rebasing?

* cosmetic fix

* correct key is sequence_parallel_enabled

3ff1a10f

Move distributed Adam unit test to contrib dir (#1406) · 57f890a7

Tim Moon authored Jun 22, 2022

* Increase default bucket size in distributed Adam

* Move distributed Adam unit test to contrib tests

Integrate into unit testing framework

* Tweak hyperparameters for dist Adam optimizer test

Improves numerical stability so we can keep tight tolerances. Adopting suggestions from @crcrpar.

* Use distributed test infrastructure in distributed Adam unit test

Suggestion from @crcrpar.

57f890a7

22 Jun, 2022 2 commits

Temporary Solution to Let `FusedAdam` support BFloat16 (#1407) · 81f8ba79

Masaki Kozuki authored Jun 22, 2022

* add temporary dispatch of double, float, half, bfloat16

* fusedadam of bfloat16

* Add bfloat16 path to FusedAdam

81f8ba79

Gradient clipping with fused kernels (#1405) · dcb02fcf

Tim Moon authored Jun 21, 2022

* Gradient clipping routine with fused kernels

Identical API as PyTorch. Falls back to PyTorch impl when not computing L2 norm.

* Add unit test for gradient clipping

* Add fp16 case to gradient clipping unit test

* Tweaks to grad clipping unit test

Review suggestions from @crcrpar

* Debug gradient clipping tests

When checking that incorrect results produce assertion errors, make sure to generate a discrepancy outside the range of numerical error.

dcb02fcf

16 Jun, 2022 1 commit

Remove legacy fuser usage from multihead attention in contrib in favor of the... · 1403c21a

Kevin Stephano authored Jun 15, 2022

Remove legacy fuser usage from multihead attention in contrib in favor of the default which should be nvfuser.  Modify test scripts to activate fusion. (#1403)

1403c21a

14 Jun, 2022 3 commits
- Merge pull request #1401 from timmoon10/dist-adam-zero · 5ffb22d0
  Thor Johnsen authored Jun 14, 2022
```
ZeRO-2 support in DistributedFusedAdam
```
  5ffb22d0
- Update documentation to reflect DistributedFusedAdam uses AdamW · 846f7f8a
  Tim Moon authored Jun 14, 2022
```
Adjust test options to have tighter tolerances.
```
  846f7f8a
- Update dist Adam test to use updated API · e2af089c
  Tim Moon authored Jun 13, 2022
  
  e2af089c