Commits · 8ea429504207523c64e6b82e41fdf90d4ffe55e4 · tsoc / openmm

23 Sep, 2024 1 commit

Optimize PME spread charge kernel (#4633) · 8ea42950

Anton Gorenko authored Sep 24, 2024

* PME_ORDER threads process one atom;
* PME_ORDER threads access consecutive addresses;
* No need to permute z indices with zindexTable;
* finishSpreadCharge is needed only with fixed point charge spreading;

8ea42950

18 Aug, 2023 1 commit

Amoeba and ATM Force minor cleanup (#4195) · f2bdcc75

bdenhollander authored Aug 18, 2023

* Amoeba minor cleanup

- Fix variable name in string
- Remove odd space between variable and period that is inconsistently styled

* Replaces random tabs with spaces in ATM Force

f2bdcc75

17 Aug, 2022 1 commit
- Improved support for devices without 64 bit atomics (#3737) · ae686364
  Peter Eastman authored Aug 17, 2022
  
  ae686364
22 Jul, 2022 1 commit

Final HIP Platform implementation for AMD GPUs on ROCm (#3338) · a39fa14a

Adel Johar authored Jul 22, 2022



* Support kernel files with extensions of any length (like .hip)

* Do not allow to replace symbols in single-line comments

* Add OPENMM_BUILD_COMMON CMake option

It allows to build and install common platform files even if
CUDA or OpenCL platforms are not built.
This is required for HIP platform (openmm-hip) if ROCm OpenCL
packages are not installed.

* Add an option for Python wrapper to install into user packages

OPENMM_PYTHON_USER_INSTALL is OFF be default.

* Support FFT backends in Amoeba plugin

The HIP platform supports FFT backends, this commit moves
findLegalFFTDimension to ComputeContext, so platforms can have their own
implementations.

* Compatibility for common platform w/ new HIP platform

* Do not use volatile with private and local AtomData parameters on HIP

The generated code is not optimal, for example, the compiler generates
flat_load instructions instead of ds_read.

* Tune launch bounds for PME grid-related kernels and add WA for RDNA

Force the compiler to use all registers for gridSpreadCharge and
gridInterpolateForce by limiting max waves per EU to 1 on CDNA GPUs,
RDNA GPUs work better without it.

* Optimize atom data structs in GBSA and Amoeba on HIP

Manually rearrange fields, add paddings and force alignments to
have faster accesses to shared memory: ds_read and ds_write may
work slower if addresses are not aligned by 16 bytes.
Co-authored-by: Anton Gorenko <anton@streamhpc.com>
Co-authored-by: Nick Curtis <nicholas.curtis@amd.com>

a39fa14a

27 Jan, 2022 1 commit
- Fixed potential invalid memory access (#3428) · 995c6318
  Peter Eastman authored Jan 26, 2022
```
* Fixed potential invalid memory access

* Fixed exception
```
  995c6318
04 Oct, 2021 1 commit

Use cuCtxPushCurrent() and cuCtxPopCurrent() for selecting CUDA context (#3258) · c456dd54

Peter Eastman authored Oct 04, 2021

* Use cuCtxPushCurrent() and cuCtxPopCurrent() for selecting CUDA context

* Fixed errors in amoeba coda

* Fixed more errors in context selection

c456dd54

22 May, 2021 1 commit

Converted AMOEBA to common platform (#3120) · 8e8923a7

Peter Eastman authored May 22, 2021

* Began converting AMOEBA to common platform

* Beginning of OpenCL platform for AMOEBA

* Converted AmoebaVdwForce to common platform

* Cleaned up reference AMOEBA tests

* Began converting AmoebaMultipoleForce to common platform

* Continue converting AmoebaMultipoleForce to common platform

* Bug fixes

* Bug fix

* Continue converting AmoebaMultipoleForce to common platform

* Converting AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce to common platform

* Converting AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce to common platform

* Creating OpenCL version of AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce

* Creating OpenCL version of AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce

* Creating OpenCL version of AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce

* Converted arrays from real3 to real

* Bug fix to OpenCL AmoebaGeneralizedKirkwoodForce

* Fixes for AMD GPUs

* Began converting HippoNonbondedForce to common platform

* Continuing to convert HippoNonbondedForce to common platform

* Continuing to convert HippoNonbondedForce to common platform

* Working on unifying PME kernels

* Fixed error on devices without 64 bit atomics

* Unified PME kernels

* Converted HippoNonbondedForce to common platform

* Creating OpenCL implementation of HippoNonbondedForce

* Continuing OpenCL implementation of HippoNonbondedForce

* Mostly finished OpenCL implementation of HippoNonbondedForce

* Eliminated three component vector types in host code

* Fix errors on CPU OpenCL

* Skip double precision tests for AMOEBA on OpenCL

* Bug fixes

* Bug fixes

* Fixed compilation error

8e8923a7