compete removal of sort engine; avoid third pass to complete collision search by slowing down the first pass
add second bucket sort pass, remove radix sort infrastructure, tune memory sizing
reorganize GPU code; begin refactoring collision search to avoid CUB's radix sort, which does too much work for collision purposes
more cleanup
remove more debug; fix batch inversion and add performance optimizations for small problems
mix hashtable-based and sort-based collision search to be more efficient across a range of problem sizes
make small-array collision search more space efficient and store overflow offsets in global memory; add large-array collision search (WIP)
fix makefile for CUDA 12