* All Cuda* objects now take in a ContextImpl* during
construction, and save the context instead of individual
handles.
* Since we no longer use the legacy default stream, we need to
explicitly synchronize the stream before performing GPU->CPU
transfers, and CudaBuffer is responsible for such synchronization
when asked to perform GPU to CPU transfers.
* Remove all manual syncs and relegate syncing to CudaBuffer
before performing GPU to CPU transfers.
Change-Id: Ic73cb24174a1e09842827323280e90241716cc20
* Added CudaSparseMatrix to manage and operate on sparse matrices with
cuSparse.
* Added tests for CudaSparseMatrix.
* Added a new sparse linear operator benchmark.
Change-Id: Id09df46de3b40be1f14441528088b54dab5844af