Commits · b1a21fd456208ffe40fefc23dd97628bb0a56440 · tsoc / openmm

02 Jun, 2026 1 commit

Enable split PME streams for HIP LJPME · b1a21fd4

one authored May 31, 2026

Run Coulomb and dispersion reciprocal PME work on separate HIP queues for
LJPME when PME streams are enabled. Use separate grids, sorters, events, and
energy buffers so the two reciprocal branches can overlap safely.

Keep the behavior HIP-only based on RTX4090 CUDA profiling, where the same
split increased PME spread/list contention and regressed apoa1ljpme.

b1a21fd4

31 May, 2026 1 commit

Revert "Add wave64 LDS spreading in HIP LJ-PME" · c26187aa

one authored May 31, 2026

This reverts commit 4e7070c2.

The split-pme-stream optimization to be implemented will put
LJ-PME reciprocal kernels on a non-critical path.

c26187aa

11 May, 2026 1 commit

Tune HIP PME kernel launch block sizes · 20e4b551

one authored May 12, 2026

Use explicit 128-thread block launches for selected HIP PME kernels that
benefit from larger blocks. Keep the platform default block size unchanged,
and leave small-system grid indexing and charge spreading on the existing
default launch configuration.

The heuristic applies 128-thread launches to finishSpreadCharge on HIP, and
uses 128-thread launches for findAtomGridIndex and gridSpreadCharge only for
larger systems. Coulomb PME and LJPME dispersion paths are handled in
parallel, while interpolation and energy evaluation remain unchanged.

20e4b551

06 May, 2026 1 commit
- Add wave64 LDS spreading in HIP LJ-PME · 4e7070c2
  one authored Apr 30, 2026
  
  4e7070c2
29 Apr, 2026 1 commit
- Avoid host wait for PME post-computation synchronization (#5271) · 0aee8050
  one authored Apr 30, 2026
```
* Avoid host wait when synchronizing PME queue

* Remove comment
```
  0aee8050
24 Apr, 2026 1 commit
- Avoid host wait in PME post-sync · c2d9cc7b
  one authored Apr 24, 2026
  
  c2d9cc7b
17 Apr, 2026 1 commit

Split LJ-PME atom-grid sorting from Coulomb PME · c1d643e2

one authored Apr 17, 2026

Avoid forcing Coulomb PME to re-sort whenever LJ-PME is enabled, and give dispersion PME its own atom-grid index and sort state so the performance impact can be measured independently.

c1d643e2

16 Apr, 2026 1 commit
- Tune sorting threshold in Coulomb PME (v3.1) · ee4ca894
  one authored Apr 16, 2026
  
  ee4ca894
10 Feb, 2026 1 commit
- Avoid error in updateParametersInContext() when there are no exceptions (#5209) · 122dbde2
  Peter Eastman authored Feb 10, 2026
  
  122dbde2
22 Oct, 2025 1 commit
- Avoid incorrectly adding plasma correction (#5117) · 895fcd4c
  Peter Eastman authored Oct 22, 2025
  
  895fcd4c
23 Sep, 2025 1 commit

Update file headers (#5074) · 05472c9a

Evan Pretti authored Sep 23, 2025

* Replace SimTK-containing file headers

* Update file headers for new Tinker reader files added

05472c9a

12 Sep, 2025 1 commit

Add constant potential method (#4870) · f55abcaa

Evan Pretti authored Sep 12, 2025



* Initial implementation of C++ API

* Add kernel interface and information for API generation

* API updates for updating electrode parameters

* Add serialization proxy for ConstantPotentialForce

* Update file headers

* Add CG error tolerance and fix units on getCharges() return value

* Initial implementation of matrix solver

* Fixes and conjugate gradient solver

* Try to fix Linux and Windows builds

* Make sure charge constraint target is on total charge

* Restore handling of exceptions like NonbondedForce since they won't involve electrode atoms

* Ameliorate numerical instability in constrained conjugate gradient

* Fix uninitialized pointers, memory leak, and style

* Set CG tolerance units in Python API

* Test ConstantPotentialForce serialization

* Read/write ExceptionsUsePeriodicBoundaryConditions as bool

* Improve constrained conjugate gradient robustness to roundoff error accumulation

* Recompute matrix if electrode atoms move due to setPositions()

* Tolerance is now in gradient (potential) units again

* Add neutralizing background correction

* Add Python API tests

* Fixes for CG and nonbonded exceptions

* Add initial tests checking against existing NonbondedForce behavior

* Expand test suite and fix some implementation issues

* Add additional tests using larger reference system

* Add Gaussian test

* Finish test against reference computation

* CPU platform implementation

* Fixes for compilation on some platforms

* Fixes for constant potential with AVX/AVX2

* Test linking CPU PME library to constant potential test directly

* Older SWIG versions don't support Python set to C++ set conversion

* Add user guide entry

* Increase speed of reference test

* Conditional building constant potential CPU test is unreliable

* Debugging

* Miscellaneous fixes and improvements for CI

* Cache charges so solver will not run if system and coordinates have not changed

* Preconditioner flag, stability, and automatic detection improvements

* Add GPU platform-specific constant potential kernel classes

* PME and device-host I/O changes to support constant potential

* Initial common constant potential implementation

* Constant potential fixes:

* Fix preconditioner PME position/charge save/restore logic

* Fix reduction synchronization in constant potential solver kernels

* Add double-float accumulation for conjugate gradient solver when
  double unsupported by hardware

* Improve conditioning of a test system, and make sure particles are in or
out of cutoff for consistency and ease of comparing between platforms

* Reorder guess charges for CG when atom reordering changes positions

* Remove PME queue for now

* Trying to debug optimized direct space derivative kernel

* Remove extraneous debugging lines

* Style updates; just make CPU preconditioner double precision

* Debugging updated optimized direct derivatives kernel for all but OpenCL CPU

* OpenCL CPU implementation of direct space derivatives, and cleanup

* Try to make test even shorter to not time out on CI

* Temporary - Debugging

* Debugging

* Debugging

* Debugging

* Debugging

* Remove debugging code and fix reduction synchronization

* Fix other reductions

* Debugging - are tests hanging or just slow on CI?

* Debugging

* Debugging

* Fix macro for case when double precision is available on hardware

* Remove changes for debugging again

* Try to improve matrix solver cache locality by uploading transpose

* Fixes for atom ordering and periodic images

* Can't rely on reorder listener for cell offset updates

* Test reducing number of contexts and timing for CI

* Debugging

* Remove timing code and revert debugging changes

* Matrix solver and plasma term optimizations

* Reduce CG solver kernel calls and downloads

* Don't read back convergence flag from global memory

* Update PME due to refactoring in master branch

* Faster matrix solver (1st step)

* Faster matrix solver for CUDA

* Faster matrix solver compatibility with non-CUDA platforms

* Matrix solver fixes

* Use warp shuffle reductions when possible

* Attempt to work around intermittent compiler crash in Intel CPU OpenCL

* Optimize CG solver kernel 1

* Rework CG solver so some kernels can use more than 1 block

* Don't run out of shared memory

* Asynchronously download convergence flag while clearing buffers

---------
Co-authored-by: Evan Pretti <pretti@sh03-17n15.int>

f55abcaa

06 Sep, 2025 1 commit
- Fixed bug in calculating plasma self energy correction with offsets (#5069) · bb3f4f82
  Peter Eastman authored Sep 06, 2025
  
  bb3f4f82
06 Aug, 2025 1 commit
- Reducing sorting overhead for small systems with PME (#5027) · fdea63e9
  Peter Eastman authored Aug 06, 2025
  
  fdea63e9
09 Jul, 2025 1 commit

Unified storage of global parameters (#5002) · d47ea1de

Peter Eastman authored Jul 09, 2025

* Unified storage of global parameters

* Fixes to CUDA and HIP

* Store global parameters as real instead of float

d47ea1de

05 May, 2025 1 commit

Common implementation of NonbondedForce (#4922) · 2443dcee

Peter Eastman authored May 05, 2025

* Use common API for kernels

* More code uses common interface

* Bug fixes

* Unified interface for sorting

* Simplified interface for FFT

* Use common event API for synchronization

* Minor changes to make code more consistent between platforms

* Common implementation of NonbondedForce

* Bug fixes

* Flag to enable list of single pairs

* CUDA and OpenCL use common implementation of NonbondedForce

* Fixed compilation error

* HIP uses common implementation of NonbondedForce

2443dcee

28 Apr, 2025 1 commit

Unified interface for queues (#4913) · dd320bcf

Peter Eastman authored Apr 28, 2025

* Unified interface for queues

* Simplified stream handling in CudaFFT3D

* HIP implementation of ComputeQueue

dd320bcf

25 Apr, 2025 1 commit

Uniform interface for FFTs (#4911) · 01e99e77

Peter Eastman authored Apr 25, 2025

* Unified interface for FFTs

* AMOEBA uses unified interface for FFTs

* HIP implementation of common FFT interface

01e99e77

23 Apr, 2025 1 commit

Add correction for self energy of neutralizing plasma (#4907) · a3909c8e

Peter Eastman authored Apr 23, 2025

* Add correction for self energy of neutralizing plasma

* Fixed compilation errors

* Update total charge in copyParametersToContext()

* Bug fixes

* Fixed compilation errors in HIP

* Bug fix

a3909c8e

13 Jan, 2025 1 commit
- Allow more changes to nonbonded exceptions (#4766) · a3628b48
  Peter Eastman authored Jan 13, 2025
  
  a3628b48
22 Nov, 2024 1 commit

updateParametersInContext() can modify parameter offsets (#4732) · ecbe32b0

Peter Eastman authored Nov 22, 2024

* updateParametersInContext() can modify parameter offsets

* Reordering respects parameter offsets

* Implemented for CUDA and HIP

ecbe32b0

23 Sep, 2024 1 commit

Optimize PME spread charge kernel (#4633) · 8ea42950

Anton Gorenko authored Sep 24, 2024

* PME_ORDER threads process one atom;
* PME_ORDER threads access consecutive addresses;
* No need to permute z indices with zindexTable;
* finishSpreadCharge is needed only with fixed point charge spreading;

8ea42950

06 Sep, 2024 1 commit

Optimize updateParametersInContext() (#4610) · 78902bed

Peter Eastman authored Sep 06, 2024

* Optimize CustomNonbondedForce.updateParametersInContext()

* Optimized uploading changed values to GPU

* Optimized updateParametersInContext() for lots of bonded forces

* Optimized updateParametersInContext() for CustomExternalForce

* Optimized updateParametersInContext() for NonbondedForce

* Code changes for HIP platform

78902bed

12 Dec, 2023 1 commit
- Merged more code into common platform (#4346) · 5739788a
  Peter Eastman authored Dec 12, 2023
```
* Common implementation of BondedUtilities

* Common implementation of UpdateStateDataKernel
```
  5739788a
23 May, 2023 1 commit

Skip neighbor list for very small systems (#4070) · 655518c3

Peter Eastman authored May 22, 2023

* Skip neighbor list for very small systems

* Fixed typos

* Don't skip box size check when not using neighbor list

* Made test larger to ensure it uses neighbor list

655518c3

14 Feb, 2023 1 commit
- Remove unused offset variables (#3961) · 611bd817
  bdenhollander authored Feb 13, 2023
```
- Appear to be copy and pasted from getPositions and were not removed
```
  611bd817
12 Sep, 2022 1 commit
- Ensure valid atom order after loading a checkpoint (#3771) · a056d5a3
  Peter Eastman authored Sep 12, 2022
  
  a056d5a3
08 Sep, 2022 1 commit
- Ensure valid atom order after loading a checkpoint · c29e203b
  peastman authored Sep 08, 2022
  
  c29e203b
17 Aug, 2022 1 commit
- Improved support for devices without 64 bit atomics (#3737) · ae686364
  Peter Eastman authored Aug 17, 2022
  
  ae686364
17 May, 2022 1 commit
- Very minor optimizations (#3602) · 109f6b25
  Peter Eastman authored May 17, 2022
  
  109f6b25
01 Mar, 2022 1 commit

Improved temperature reporting for Drude particles (#3486) · a5e42f57

Peter Eastman authored Mar 01, 2022

* DrudeLangevinIntegrator has getSystemTemperature()

* DrudeNoseHooverIntegrator has getSystemTemperature()

* StateDataReporter reports system temperature for Drude systems

* Fixed incorrect return type

a5e42f57

10 Jan, 2022 1 commit
- Fixed uninitialized memory access (#3399) · 6eabb453
  Peter Eastman authored Jan 10, 2022
  
  6eabb453
20 Nov, 2021 1 commit
- Fixed errors when running on multiple GPUs (#3340) · ed9df876
  Peter Eastman authored Nov 19, 2021
  
  ed9df876
16 Sep, 2021 1 commit
- Allow querying current step count (#3248) · 22da37a4
  Peter Eastman authored Sep 16, 2021
```
* Allow querying current step count

* Fixed error building Python wrapper
```
  22da37a4
15 Jul, 2021 1 commit
- Allow disabling the direct space calculation in NonbondedForce (#3173) · 1344f2e0
  Peter Eastman authored Jul 15, 2021
  
  1344f2e0
15 Jun, 2021 1 commit
- Disable PME stream when using CPU PME (#3148) · a29596c7
  Peter Eastman authored Jun 15, 2021
  
  a29596c7
22 May, 2021 1 commit

Converted AMOEBA to common platform (#3120) · 8e8923a7

Peter Eastman authored May 22, 2021

* Began converting AMOEBA to common platform

* Beginning of OpenCL platform for AMOEBA

* Converted AmoebaVdwForce to common platform

* Cleaned up reference AMOEBA tests

* Began converting AmoebaMultipoleForce to common platform

* Continue converting AmoebaMultipoleForce to common platform

* Bug fixes

* Bug fix

* Continue converting AmoebaMultipoleForce to common platform

* Converting AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce to common platform

* Converting AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce to common platform

* Creating OpenCL version of AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce

* Creating OpenCL version of AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce

* Creating OpenCL version of AmoebaMultipoleForce and AmoebaGeneralizedKirkwoodForce

* Converted arrays from real3 to real

* Bug fix to OpenCL AmoebaGeneralizedKirkwoodForce

* Fixes for AMD GPUs

* Began converting HippoNonbondedForce to common platform

* Continuing to convert HippoNonbondedForce to common platform

* Continuing to convert HippoNonbondedForce to common platform

* Working on unifying PME kernels

* Fixed error on devices without 64 bit atomics

* Unified PME kernels

* Converted HippoNonbondedForce to common platform

* Creating OpenCL implementation of HippoNonbondedForce

* Continuing OpenCL implementation of HippoNonbondedForce

* Mostly finished OpenCL implementation of HippoNonbondedForce

* Eliminated three component vector types in host code

* Fix errors on CPU OpenCL

* Skip double precision tests for AMOEBA on OpenCL

* Bug fixes

* Bug fixes

* Fixed compilation error

8e8923a7

19 Mar, 2021 1 commit
- Converted more code to common platform (#3073) · 98d81730
  Peter Eastman authored Mar 19, 2021
```
* Converted more code to common platform

* Converted more code to common platform
```
  98d81730
27 Dec, 2020 1 commit
- Updated to latest version of OpenCL C++ API (#2962) · ee037398
  peastman authored Dec 27, 2020
  
  ee037398
23 Dec, 2020 1 commit
- Fixed potential out of range index (#2960) · 84132b19
  peastman authored Dec 23, 2020
  
  84132b19