Commits · 609260cd2db52abc4eabc70521442aa31901bbbc · ModelZoo / ResNet50_tensorflow

21 Jun, 2019 1 commit

NCF XLA and Eager tests with a refactor of resnet flags to make this cleaner. (#7067) · a68f65f8

Toby Boyd authored Jun 21, 2019

* XLA FP32 and first test

* More XLA benchmarks FP32.

* Add eager to NCF and refactor resnet.

* fix v2_0 calls and more flag refactor.

* Remove extra flag args.

* 90 epoch default

* add return

* remove xla not used by estimator.

* Remove duplicate run_eagerly.

* fix flag defaults.

* Remove fp16_implementation flag option.

* Remove stop early on mlperf test.

* remove unneeded args.

* load flags from keras mains.

a68f65f8

13 Jun, 2019 2 commits
- fix ctl case; add check for 2.0 · f6f04066
  guptapriya authored Jun 11, 2019
  
  f6f04066
- Clean up unused flags etc · 71c6a697
  guptapriya authored Jun 07, 2019
  
  71c6a697
03 Jun, 2019 1 commit
- remove cloning flag · 9511801a
  guptapriya authored Jun 03, 2019
  
  9511801a
31 May, 2019 2 commits
- Fix internal lint errors (#6937) · 7546a9e3
  Haoyu Zhang authored May 31, 2019
  
  7546a9e3
- Fix various lint errors (#6934) · ba415414
  Haoyu Zhang authored May 31, 2019
```
* Fix various lint errors

* Fix logging format
```
  ba415414
29 May, 2019 1 commit

Add flag to use custom training loop for keras NCF model. (#6905) · b5a69819

Bruce Fontaine authored May 28, 2019

* Add flag to use custom training loop for keras NCF model.

* Add error check to NCF model for custom training loop + tf1.0.

b5a69819

28 May, 2019 1 commit

Add a custom training loop for NCF model with TF2.0 (#6899) · 4c1d95cc

Bruce Fontaine authored May 28, 2019

* Add a custom training loop for NCF model with TF2.0

* Fix long line in ncf_keras_main.py

* Remove dataset repeat when using custom training loop.

4c1d95cc

24 May, 2019 1 commit

Add early stopping logic to ncf keras when desired threshold is met. Also... · 7033c8a2

Priya Gupta authored May 23, 2019

Add early stopping logic to ncf keras when desired threshold is met. Also change the default batch size to match the tuned hyperparams

7033c8a2

29 Apr, 2019 1 commit

Add benchmarks with the --cloning flag to Resnet and NFC. (#6675) · af47736d

Igor authored Apr 29, 2019

* Add benchmarks with the --cloning flag to Resnet and NFC.

* Renamed cloning to clone_model_in_keras_dist_strat. Dropped a few tests that aren't essential.

* Fixed up the formatting after re-naming the flag to a much longer  name.  Thanks, lint.
* Fixed the lint error in nfc_common.py

af47736d

20 Apr, 2019 1 commit

Remove contrib imports, or move them inline (#6591) · 8ff9eb54

Shining Sun authored Apr 19, 2019

* Remove contrib imports, or move them inline

* Use exposed API for FixedLenFeature

* Replace tf.logging with absl logging

* Change GFile to v2 APIs

* replace tf.logging with absl loggin in movielens

* Fixing an import bug

* Change gfile to v2 APIs in code

* Swap to keras optimizer v2

* Bug fix for optimizer

* Change tf.log to tf.keras.backend.log

* Change the loss function to keras loss

* convert another loss to keras loss

* Resolve comments and fix lint

* Add a doc string

* Fix existing tests and add new tests for DS

* Added tests for multi-replica

* Fix lint

* resolve comments

* make estimator run in tf2.0

* use compat v1 loss

* fix lint issue

8ff9eb54

01 Mar, 2019 1 commit

Keras-fy NCF Model (#6092) · 048e5bff

Shining Sun authored Mar 01, 2019

* tmp commit

* tmp commit

* first attempt (without eval)

* Bug fixes

* bug fixes

* training done

* Loss NAN, no eval

* Loss weight problem solved

* resolve the NAN loss problem

* Problem solved. Clean up needed

* Added a todo

* Remove debug prints

* Extract get_optimizer to ncf_common

* Move metrics computation back to neumf; use DS.scope api

* Extract DS.scope code to utils

* lint fixes

* Move obtaining DS above producer.start to avoid race condition

* move pt 1

* move pt 2

* Update the run script

* Wrap keras_model related code into functions

* Update the doc for softmax_logitfy and change the method name

* Resolve PR comments

* working version with: eager, DS, batch and no masks

* Remove git conflict indicator

* move reshape to neumf_model

* working version, not converge

* converged

* fix a test

* more lint fix

* more lint fix

* more lint fixes

* more lint fix

* Removed unused imports

* fix test

* dummy commit for kicking of checks

* fix lint issue

* dummy input to kick off checks

* dummy input to kick off checks

* add collective to dist strat

* addressed review comments

* add a doc string

048e5bff

08 Jan, 2019 1 commit
- update call to TPUStrategy · c8be4828
  Taylor Robie authored Jan 08, 2019
  
  c8be4828
07 Jan, 2019 6 commits

remove references to hash pipeline · d569b531
Taylor Robie authored Dec 27, 2018

d569b531

Add bisection based producer for increased scalability, enable fully... · 4fb325da

Taylor Robie authored Dec 27, 2018

Add bisection based producer for increased scalability, enable fully deterministic data production, and use the materialized and bisection producer to check each other (via expected output md5's)

4fb325da

change eval_batch_size flag from a string to an int · 1048ffd5
Taylor Robie authored Dec 26, 2018

1048ffd5
address PR comments · ec0d43ba
Taylor Robie authored Dec 21, 2018

ec0d43ba
remove 'deterministic' · 9d42f797
Taylor Robie authored Dec 21, 2018

9d42f797

rough pass at carving out existing NCF pipeline · c5ff4ec7

Taylor Robie authored Nov 18, 2018

2nd half of rough replacement pass

fix dataset map functions

reduce bias in sample selection

cache pandas work on a daily basis

cleanup and fix batch check for multi gpu

multi device fix

fix treatment of eval data padding

print data producer

replace epoch overlap with padding and masking

move type and shape info into the producer class and update run.sh with larger batch size hyperparams

remove xla for multi GPU

more cleanup

remove model runner altogether

bug fixes

address subtle pipeline hang and improve producer __repr__

fix crash

fix assert

use popen_helper to create pools

add StreamingFilesDataset and abstract data storage to a separate class

bug fix

fix wait bug and add manual stack trace print

more bug fixes and refactor valid point mask to work with TPU sharding

misc bug fixes and adjust dtypes

address crash from decoding bools

fix remaining dtypes and change record writer pattern since it does not append

fix synthetic data

use TPUStrategy instead of TPUEstimator

minor tweaks around moving to TPUStrategy

cleanup some old code

delint and simplify permutation generation

remove low level tf layer definition, use single table with slice for keras, and misc fixes

missed minor point on removing tf layer definition

fix several bugs from recombinging layer definitions

delint and add docstrings

Update ncf_test.py. Section for identical inputs and different outputs was removed.

update data test to run against the new producer class

c5ff4ec7

03 Nov, 2018 1 commit

Have async process end when all data is written. (#5652) · 424fe9f6

Reed authored Nov 02, 2018

I've noticed sometimes the async process's pool processes do not die when ncf_main.py ends and kills the async process. This commit fixes the issue.

424fe9f6

01 Nov, 2018 1 commit
- Add --use_while_loop option. (#5653) · 826eea75
  Reed authored Nov 01, 2018
  
  826eea75
30 Oct, 2018 1 commit

Merges TPU-TC optimizations into HEAD. (#5635) · b8318fd3

Tayo Oguntebi authored Oct 29, 2018

* Merges TPU-TC optimizations into HEAD.

* Split a line that went over 80 from a tab.

* Remove trailing whitespace.

b8318fd3

29 Oct, 2018 1 commit
- Add option to not use estimator. (#5623) · 0c0860ed
  Reed authored Oct 29, 2018
```
The option is --nouse_estimator
```
  0c0860ed
26 Oct, 2018 1 commit

Split --ml_perf into two flags. (#5615) · 4298c3a3

Reed authored Oct 26, 2018

--ml_perf now just changes the model to make it MLPerf compliant. --output_ml_perf_compliance_logging adds the MLPerf compliance logs.

4298c3a3

25 Oct, 2018 1 commit

Fix crash when --ml_perf flag is not specified. (#5610) · 48a4b443

Reed authored Oct 25, 2018

The error message was:

absl.flags._exceptions.IllegalFlagValueError: flag --ml_perf=None: ('Non-boolean argument to boolean flag', 'None')

48a4b443

24 Oct, 2018 1 commit

Add logging calls to NCF (#5576) · 780f5265

Taylor Robie authored Oct 24, 2018

* first pass at __getattr__ abuse logger

* first pass at adding tags to NCF

* minor formatting updates

* fix tag name

* convert metrics to python floats

* getting closer...

* direct mlperf logs to a file

* small tweaks and add stitching

* update tags

* fix tag and add a sudo call

* tweak format of run.sh

* delint

* use distribution strategies for evaluation

* address PR comments

* delint and fix test

* adjust flag validation for xla

* add prefix to distinguish log stitching

* fix index bug

* fix clear cache for root user

* dockerize cache drop

* TIL some regex magic

780f5265

20 Oct, 2018 1 commit
- Add XLA support to NCF (#5572) · f2b702a0
  Reed authored Oct 19, 2018
  
  f2b702a0
18 Oct, 2018 1 commit

Reorder NCF data pipeline (#5536) · 19d4eaaf

Taylor Robie authored Oct 18, 2018

* intermediate commit

finish replacing spillover with resampled padding

intermediate commit

* resolve merge conflict

* intermediate commit

* further consolidate the data pipeline

* complete first pass at data pipeline refactor

* remove some leftover code

* fix test

* remove resampling, and move train padding logic into neumf.py

* small tweaks

* fix weight bug

* address PR comments

* fix dict zip. (Reed led me astray)

* delint

* make data test deterministic and delint

* Reed didn't lead me astray. I just can't read.

* more delinting

* even more delinting

* use resampling for last batch padding

* pad last batch with unique data

* Revert "pad last batch with unique data"

This reverts commit cbdf46efcd5c7907038a24105b88d38e7f1d6da2.

* move padded batch to the beginning

* delint

* fix step check for synthetic data

19d4eaaf

14 Oct, 2018 1 commit
- Make flagfile sharing robust to distributed filesystems and multi-worker setups. (#5521) · 91b2debd
  Taylor Robie authored Oct 14, 2018
```
* move flagfile into the cache_dir

* remove duplicate code

* delint
```
  91b2debd
11 Oct, 2018 1 commit
- Added option to use_subprocess or not in ncf_main.py. · d4ac494f
  Shawn Wang authored Oct 11, 2018
  
  d4ac494f
10 Oct, 2018 1 commit

Add --use_synthetic_data option to NCF. (#5468) · 75d592e9

Reed authored Oct 10, 2018

* Add --use_synthetic_data option to NCF.

* Add comment to _SYNTHETIC_BATCHES_PER_EPOCH

* Fix test

* Hopefully fix lint issue

75d592e9

05 Oct, 2018 1 commit

Fix/ncf eval default (#5438) · aec1fec6

Taylor Robie authored Oct 04, 2018

* improve default handling for eval_batch_size

* return eval_batch_size default to None

* fix syntax error

aec1fec6

03 Oct, 2018 1 commit

Move evaluation to .evaluate() (#5413) · c494582f

Taylor Robie authored Oct 02, 2018

* move evaluation from numpy to tensorflow

fix syntax error

don't use sigmoid to convert logits. there is too much precision loss.

WIP: add logit metrics

continue refactor of NCF evaluation

fix syntax error

fix bugs in eval loss calculation

fix eval loss reweighting

remove numpy based metric calculations

fix logging hooks

fix sigmoid to softmax bug

fix comment

catch rare PIPE error and address some PR comments

* fix metric test and address PR comments

* delint and fix python2

* fix test and address PR comments

* extend eval to TPUs

c494582f

02 Oct, 2018 1 commit
- Add flags for adam hyperparameters (#5428) · f3be93a7
  Reed authored Oct 02, 2018
  
  f3be93a7
20 Sep, 2018 1 commit

Fix/ncf mlperf tweaks: robustness and determinism (#5334) · 4dc1080d

Taylor Robie authored Sep 19, 2018

* bug fixes and add seed

* more random corrections

* make cleanup more robust

* return cleanup fn

* delint and address PR comments.

* delint and fix tests

* delinting is never done

* add pipeline hashing

* delint

4dc1080d

22 Aug, 2018 1 commit

Fix convergence issues for MLPerf. (#5161) · 64710c05

Reed authored Aug 22, 2018

* Fix convergence issues for MLPerf.

Thank you to @robieta for helping me find these issues, and for providng an algorithm for the `get_hit_rate_and_ndcg_mlperf` function.

This change causes every forked process to set a new seed, so that forked processes do not generate the same set of random numbers. This improves evaluation hit rates.

Additionally, it adds a flag, --ml_perf, that makes further changes so that the evaluation hit rate can match the MLPerf reference implementation.

I ran 4 times with --ml_perf and 4 times without. Without --ml_perf, the highest hit rates achieved by each run were 0.6278, 0.6287, 0.6289, and 0.6241. With --ml_perf, the highest hit rates were 0.6353, 0.6356, 0.6367, and 0.6353.

* fix lint error

* Fix failing test

* Address @robieta's feedback

* Address more feedback

64710c05

31 Jul, 2018 1 commit
- Fix crash when --eval_batch_size is not set. (#4955) · e4034bef
  Reed authored Jul 31, 2018
  
  e4034bef
30 Jul, 2018 1 commit

NCF pipeline refactor (take 2) and initial TPU port. (#4935) · 6518c1c7

Taylor Robie authored Jul 30, 2018

* intermediate commit

* ncf now working

* reorder pipeline

* allow batched decode for file backed dataset

* fix bug

* more tweaks

* parallize false negative generation

* shared pool hack

* workers ignore sigint

* intermediate commit

* simplify buffer backed dataset creation to fixed length record approach only. (more cleanup needed)

* more tweaks

* simplify pipeline

* fix misplaced cleanup() calls. (validation works\!)

* more tweaks

* sixify memoryview usage

* more sixification

* fix bug

* add future imports

* break up training input pipeline

* more pipeline tuning

* first pass at moving negative generation to async

* refactor async pipeline to use files instead of ipc

* refactor async pipeline

* move expansion and concatenation from reduce worker to generation workers

* abandon complete async due to interactions with the tensorflow threadpool

* cleanup

* remove per...

6518c1c7

20 Jun, 2018 1 commit

Wide Deep refactor and deep movies (#4506) · 20070ca4

Taylor Robie authored Jun 20, 2018

* begin branch

* finish download script

* rename download to dataset

* intermediate commit

* intermediate commit

* misc tweaks

* intermediate commit

* intermediate commit

* intermediate commit

* delint and update census test.

* add movie tests

* delint

* fix py2 issue

* address PR comments

* intermediate commit

* intermediate commit

* intermediate commit

* finish wide deep transition to vanilla movielens

* delint

* intermediate commit

* intermediate commit

* intermediate commit

* intermediate commit

* fix import

* add default ncf csv construction

* change default on download_if_missing

* shard and vectorize example serialization

* fix import

* update ncf data unittests

* delint

* delint

* more delinting

* fix wide-deep movielens serialization

* address PR comments

* add file_io tests

* investigate wide-deep test failure

* remove hard coded path and properly use flags.

* address file_io test PR comments

* missed a hash_bucked_size

20070ca4

12 Jun, 2018 1 commit

Transformer multi gpu, remove multi_gpu flag, distribution helper functions (#4457) · 29c9f985

Katherine Wu authored Jun 12, 2018

* Add DistributionStrategy to transformer model

* add num_gpu flag

* Calculate per device batch size for transformer

* remove reference to flags_core

* Add synthetic data option to transformer

* fix typo

* add import back in

* Use hierarchical copy

* address PR comments

* lint

* fix spaces

* group train op together to fix single GPU error

* Fix translate bug (sorted_keys is a dict, not a list)

* Change params to a default dict (translate.py was throwing errors because params didn't have the TPU parameters.)

* Address PR comments. Removed multi gpu flag + more

* fix lint

* fix more lints

* add todo for Synthetic dataset

* Update docs

29c9f985