1. Add a version history
2. Update copyright years across the code base
3. Run format_all.sh
4. Update version strings from 2.1.0 to 2.2.0 in the docs and
elsewhere.
Change-Id: I46d8d479d54bd6002d532785e67342106e73c9ac
Parallel implementations for right-multiply by dense vector for:
- Partitioned matrix view
- Block-sparse matrix
- CRS matrix (non-symmetric only)
When coupled with non-interleaving indexes in parallel for, this
simple aproach provides a reasonable speedup.
For example, in CRS case difference with GPGPU approach reduces
closer to memory throughput ratio for high enough core count.
./bin/spmv_benchmark
-------------------------------------------------------------------
Benchmark Time
-------------------------------------------------------------------
BM_BlockSparseRightMultiplyAndAccumulateBA/1 28.5 ms
BM_BlockSparseRightMultiplyAndAccumulateBA/2 15.7 ms
BM_BlockSparseRightMultiplyAndAccumulateBA/4 9.01 ms
BM_BlockSparseRightMultiplyAndAccumulateBA/8 5.60 ms
BM_BlockSparseRightMultiplyAndAccumulateBA/16 3.86 ms
BM_BlockSparseRightMultiplyAndAccumulateBA/28 3.84 ms
BM_BlockSparseRightMultiplyAndAccumulateUnstructured/1 23.8 ms
BM_BlockSparseRightMultiplyAndAccumulateUnstructured/2 15.0 ms
BM_BlockSparseRightMultiplyAndAccumulateUnstructured/4 8.01 ms
BM_BlockSparseRightMultiplyAndAccumulateUnstructured/8 4.02 ms
BM_BlockSparseRightMultiplyAndAccumulateUnstructured/16 2.39 ms
BM_BlockSparseRightMultiplyAndAccumulateUnstructured/28 1.68 ms
BM_BlockSparseLeftMultiplyAndAccumulateBA 30.7 ms
BM_BlockSparseLeftMultiplyAndAccumulateUnstructured 41.5 ms
BM_CRSRightMultiplyAndAccumulateBA/1 24.1 ms
BM_CRSRightMultiplyAndAccumulateBA/2 13.6 ms
BM_CRSRightMultiplyAndAccumulateBA/4 8.70 ms
BM_CRSRightMultiplyAndAccumulateBA/8 5.34 ms
BM_CRSRightMultiplyAndAccumulateBA/16 3.99 ms
BM_CRSRightMultiplyAndAccumulateBA/28 4.00 ms
BM_CRSRightMultiplyAndAccumulateUnstructured/1 21.1 ms
BM_CRSRightMultiplyAndAccumulateUnstructured/2 10.83 ms
BM_CRSRightMultiplyAndAccumulateUnstructured/4 5.88 ms
BM_CRSRightMultiplyAndAccumulateUnstructured/8 3.68 ms
BM_CRSRightMultiplyAndAccumulateUnstructured/16 2.21 ms
BM_CRSRightMultiplyAndAccumulateUnstructured/28 1.71 ms
BM_CRSLeftMultiplyAndAccumulateBA 23.6 ms
BM_CRSLeftMultiplyAndAccumulateUnstructured 22.5 ms
BM_CudaRightMultiplyAndAccumulateBA 0.679 ms
BM_CudaRightMultiplyAndAccumulateUnstructured 0.480 ms
BM_CudaLeftMultiplyAndAccumulateBA 0.774 ms
BM_CudaLeftMultiplyAndAccumulateUnstructured 0.361 ms
./bin/partitioned_matrix_view_benchmark
-----------------------------------------------------------------
Benchmark Time
-----------------------------------------------------------------
BM_PatitionedViewRightMultiplyAndAccumulateE_Static/1 18.5 ms
BM_PatitionedViewRightMultiplyAndAccumulateE_Static/2 10.7 ms
BM_PatitionedViewRightMultiplyAndAccumulateE_Static/4 6.34 ms
BM_PatitionedViewRightMultiplyAndAccumulateE_Static/8 4.26 ms
BM_PatitionedViewRightMultiplyAndAccumulateE_Static/16 3.86 ms
BM_PatitionedViewRightMultiplyAndAccumulateE_Static/28 3.75 ms
BM_PatitionedViewRightMultiplyAndAccumulateF_Static/1 18.8 ms
BM_PatitionedViewRightMultiplyAndAccumulateF_Static/2 11.9 ms
BM_PatitionedViewRightMultiplyAndAccumulateF_Static/4 6.94 ms
BM_PatitionedViewRightMultiplyAndAccumulateF_Static/8 4.41 ms
BM_PatitionedViewRightMultiplyAndAccumulateF_Static/16 3.63 ms
BM_PatitionedViewRightMultiplyAndAccumulateF_Static/28 3.86 ms
Timings correspond to intel 8176 cpu and 2080ti nvidia gpu,
with OpenMP threading backend.
Change-Id: Idc07d0563103d057ca3c8412de81a7823fe232af
These methods were historically poorly named and every time I read code
I get confused whether they are just multiplying or multiplying and
adding. Clarifying them also gives us the changce to introduce
RightMultiply and LeftMultiply methods in the base class which will
simplify a number call sites in a subsequent CL.
Fixes https://github.com/ceres-solver/ceres-solver/issues/855
Change-Id: Ice4fb483f1acd02527a6dd753ef0c5a66037f4b0
1. Convert it from a class to a template function. Where the
template parameter is "DenseVectorType". This allows us
to have a single implementation of Conjugate Gradients
without worrying about where the matrix and the vectors
are stored or what their internal representation is.
For the case of CPU based vectors, we abstract operations
on Eigen vectors using eigen_vector_ops.
2. Introduce ConjugateGradientsLinearOperator which is
templated on DenseVectorType. It is the matrix vector
multiplication abstraction.
3. Port the tests and all usages of ConjugateGradientsSolver
to this new implementation.
4. Introduce Eigen::Vector based RightMultiply and LeftMultiply
methods into LinearOperator which by default delete to the
bare pointer based interfaces.
5. Add an identity preconditioner.
These changes are being made in preparation for adding a CUDA
based CGNR solver.
Change-Id: I9da36dc6c131856dd1a4aa7e645aaf12d25dd79b
Currently, the logic for exporting symbols is rather complicated: when
tests are enabled internal symbols are exported in addition to the
public symbols. Such logic causes several problems. (1) Test binaries
link against a Ceres build that is different from the final release
since fewer optimizations are applied if more symbols are exported. (2)
Also, some toolchains hide symbols by default breaking the existing
logic eventually causing linker errors.
Since internal symbols are not intended to be used outside of the
project, we can compile them into object files and use exactly the same
binary code both for the final build and the tests without relying on
conditionals.
By default, all symbols are now hidden unless annotated as public.
Internal symbols are explicitly marked as not being exported in case
users chose not to hide symbols by default.
Change-Id: I589dd10be2f6f438508783cf99d141af0120057b
Since Ceres is moving to using GitHub for issues, and the Google
Code URL in the current copyright header will soon become invalid,
update all the headers.
Change-Id: I1fce70375d1bcf098591f07b4d8f01a5c1e0789c