- 07 Dec, 2018 1 commit
-
-
SparkSnail authored
1.Support pytorch-operator 2.remove unsupported operator
-
- 05 Dec, 2018 1 commit
-
-
fishyds authored
* Remove unused kubernetesServer config entry in config file and schema validation
-
- 30 Nov, 2018 1 commit
-
-
fishyds authored
* [Kubeflow training service] fix bug that wrongly split kube delete cmd into 2 lines * Adjust white space
-
- 29 Nov, 2018 1 commit
-
-
fishyds authored
* Add codeDir file count validation for setClusterConfig * fix a small bug if find command is not installed * Remove codeDir validation for local training service * Remove useless import
-
- 28 Nov, 2018 1 commit
-
-
SparkSnail authored
Support aks of kuberflow training service Support nnictl set nniManagerIp
-
- 25 Nov, 2018 1 commit
-
-
QuanluZhang authored
* add one more trial job status, EARLY_STOPPED * fix datastore/nnimanager/mockeddatastore. test/webui/metrics_reader not done. USER_TO_CANCEL * fix bug * modifications based on Deshui's comments * fix bug * fix bug in remote mode
-
- 23 Nov, 2018 3 commits
-
-
SparkSnail authored
Add nniManager Ip in nnictl, pai TrainingService and kubeflow TrainingService. If users set nniManagerIp, pai and kubeflow will use this ip instead of using getIPV4() function. Web UI will also use this nniManagerIp.
-
fishyds authored
* Adjust sleep position for sdk_test.py * Exit dispather process if receive Terminate command * Add comment for sleep change in sdk_test.py
-
fishyds authored
* Use different output folder for ps and worker * Add cuda_visible_devices env var if gpuNum is 0
-
- 22 Nov, 2018 1 commit
-
-
fishyds authored
[Kubeflow training service] Update kubeflow exp job config schema to support distributed training (#387) * Support distributed training on tf-operator, for worker and ps * Update validation rule for kubeflow config * small code refactor adjustment for private methods * Use different output folder for ps and worker
-
- 20 Nov, 2018 1 commit
-
-
fishyds authored
* Kubeflow TrainingService support, v1 (#373) 1. Create new Training Service: kubeflow trainning service, use 'kubectl' and kubeflow tfjobs CRD to submit and manage jobs 2. Update nni python SDK to support new kubeflow platform 3. Update nni python SDK's get_sequende_id() implementation, read NNI_TRIAL_SEQ_ID env variable, instead of reading .nni/sequence_id file 4. This version only supports Tensorflow operator. Will add more operators' support in future versions
-