Change the loop structure of IterativeRefiner to
unconditionally refine for max_num_iterations.
This is done for two reasons.
1. We expect to use this refinement for a small number of iterations
where the convergence test is useless.
2. Eliminating the convergence test means we can restructure the loop
and save on a sparse matrix-vector multiply, saving precious
compute.
Change-Id: I6347f453a5d19d234af2a2eb1bce811048963e06
1. Add Solver::Options::use_mixed_precision_solves,
and Solver::Options::max_num_refinement_iterations.
2. Make SparseCholesky::Create return a unique_ptr.
3. SparseCholesky::Create now takes LinearSolver::Options
as an argument.
4. IterativeRefiner's constructor does not require num_cols
as an argument.
5. SparseNormalCholeskySolver now uses a separate rhs vector.
This basic implementation results in a 10% reduction in solver time
and 30% reduction in linear solver memory usage.
Change-Id: I6830f32cae2febf082d2733262eb2c9f0482b0ea
1. Add SuiteSparse::CreateDenseVectorView
2. Replace calls to SuiteSparse::CreateDenseVector with
SuiteSparse::CreateDenseVectorView.
2. Replace NULL with nullptr in suitesparse.cc and
dynamic_sparse_normal_cholesky_solver.cc
Change-Id: I94355c1dc27789e5b987a7b2850e9db6176a0914
1. Remove defines which are not used anymore.
2. Enable CXX11 threading.
3. Enable EIGEN_SPARSE by default.
Change-Id: I841ba517367e6b204475e5255c91313f01a5bcdb
With the addition of C++11 support we can simplify the parallel for code by
removing the ifdef branching. Converts coordinate_descent_minimizer.cc to use
the thread_id ParallelFor API.
Tested by building with OpenMP, C++11 threads, TBB, and no threads. All tests
pass.
Also compared timing via the bundle adjuster.
./bin/bundle_adjuster --input=../problem-744-543562-pre.txt
With OpenMP num_threads=8
Head:
Time (in seconds):
Residual only evaluation 0.807753 (5)
Jacobian & residual evaluation 4.489404 (6)
Linear solver 41.826481 (5)
Minimizer 50.745857
Total 73.294424
CL:
Time (in seconds):
Residual only evaluation 0.970483 (5)
Jacobian & residual evaluation 4.647438 (6)
Linear solver 41.781892 (5)
Minimizer 50.848904
Total 73.089983
With OpenMP num_threads=1
HEAD:
Time (in seconds):
Residual only evaluation 2.990246 (5)
Jacobian & residual evaluation 14.132090 (6)
Linear solver 79.631951 (5)
Minimizer 100.281847
Total 122.946267
CL:
Time (in seconds):
Residual only evaluation 3.075178 (5)
Jacobian & residual evaluation 13.966451 (6)
Linear solver 77.005441 (5)
Minimizer 97.568712
Total 120.410454
Change-Id: I1857d7943073be7465b6c6476bf46ab11c5475a3
1. Add explicit tests for MatrixMatrixMultiplyNaive and
MatrixTransposeMatrixMultiplyNaive
2. Add tests that exercise a variety of matrix sizes for
MatrixVectorMultiply and MatrixTransposeVectorMultiply.
Change-Id: I0b25ec346b719f19b2067848f9d4bb64c9848750
Migrate all Option and Summary structs to use
inline member initialization syntax.
This reduces the amount of code, and collocates the
default values with the documentation for the corresponding
member variable.
Change-Id: I8e6b9ee3b31464699d678667f6166ace5fc137c9
- Previously we were requesting a step-size which minimised the
polynomial in: [previous.x, current.x * factor].
- This is incorrect c/f Nocedal & Wright p60, the bounds for the
minimising step-size should be: [current.x, current.x * factor].
- However, Nocedal & Wright's bounds are insufficient when the function
can return invalid values, which we support.
- In the case that f(current) is invalid, we need to contract the
step-size to lie within: [previous.x, current.x) given that we know
that previous.x is valid so a valid step must exist within that range.
Change-Id: I67b5aaa09cc6d54cf5f264e2cf894ddc2af3f3ad
- In practice, this was not causing any errors, as
GFLAGS_INCLUDE_DIR_HINTS is set from GFLAGS_INCLUDE_DIRS so the two
would have been equivalent, but we should nonetheless be using the
one set after the find_package() call.
- Note that this would only have been active for older gflags versions
that were not built with CMake.
Change-Id: Ie54145f522c974b7685052aafb1d2f6055d22337
1. Replace HashMap and HashSet with std::unordered_map and
std::unordered_set respectively.
2. Extract the pair hasher into a struct pair_hash.
3. Delete collections_port.h
4. Convert explicit iterator based loops to auto based
loops where sensible.
Change-Id: Ib88bcd13a7463d18435639d3b771abaa52080efb
- Removes CXX11 option, and all associate paraphernalia. Ceres now
requires a compiler with full >= C++11 support. In MSVC terms this
means >= 2013 Release 4.
- This deprecates the use of CERES_STD_UNORDERED_MAP and CERES_USE_CXX11
as they will now always be defined. They will be removed from the
source in a future CL.
- For clients with CMake >= 3.8 we propagate via the exported/installed
Ceres target the CXX version that was specified when Ceres was built.
For versions < 3.8 (but >= 3.5) we specify the CXX features currently
used in the Ceres public API.
Change-Id: I535b545b10156e4426659c270a4a0649e071df0e
Even though we added support for storing the upper and
lower triangular parts of symmetric matrices in
CompressedRowSparseMatrix. RightMultiply, LeftMultiply
and SquaredColumnNorm were not modified to account for this.
This CL changes their implementation and adds thorough
tests.
Also methods that cannot work correctly with symmetric
storage now CHECK and fail, because that indicates
programmer error.
Change-Id: I76288472c8bac98db7376a79bdb6259e346ef2b7
- Testing for C++11, and subsequent testing for components using C++11
requires adding -std=c++11 to CMAKE_REQUIRED_FLAGS used by the
check_cxx_source_compiles() (and similar) macros if required by the
compiler.
- This variable is also used for the C equivalent macros, which are used
by find_package(BLAS), and would cause the detection to fail if
present.
- Thus if CXX11 was forced ON via -D (not via the GUI) then detection
would fail (with the GUI an initial configure without CXX11 would
initialise the variables and no errors would occur in later
reconfigures).
Change-Id: I08d300baf3730e926fb8e853a384badcceefeaf5
- Use cmake_dependent_option() for all threading models and encode their
respective mutual exclusivity in the dependent-option s/t the job of
ensuring mutual exclusivity is performed by CMake.
- Use of cmake_dependent_option() has the downside that options are
removed from the GUI if their requirements for availability are not
met. So if CXX11 is OFF, then neither TBB nor CXX11_THREADS will
even appear as options (they will be undefined, assumed OFF).
- Perform checks for availability of C++11 if CXX11 is enabled prior to
testing other options to handle dependence of TBB & CXX11_THREADS on
CXX11 actually being present.
- Add check for availability of C++11-specific math functions used in
jet.h before reporting C++11 as found.
- Add check for <atomic> before allowing TBB or CXX11_THREADS, as both
options require it and it may not always be available even if the
other required C++11 components are.
Change-Id: Ife8cd182aeb1978539ce2c9966910ae59ebd9177
This adds a callback mechanism to for users to get notified just
before jacobian and residual evaluations. This will enable
aggressive caching and sharing of compute between cost functions.
Change-Id: I67993726920218edf71ab9ae70c34c204756c71a
1. When Solver::Options::update_state_every_iteration = true,
the StateUpdatingCallback only updates the state on a
successful iteration. This works fine but will not work
if the user provides an EvaluationCallback, because
every call to Evaluator::Evaluate will change the state
visible to the user. This means that on unsuccessful
iterations when the user's IterationCallback is called
the state visible to the user would be the last evaluation
which did not lead to an improved cost. This CL forces
the StateUpdatingCallback to unconditionally update
the user visible state.
2. When the minimizer terminates, we update the user visible
state only if the solution is usable, otherwise the user's
state will remain the same. To maintain this invariant
we will now cache the user's state and make sure we
use it to reset it upon return if the solver is not
successful.
Change-Id: Ic1f8fa6cb10d2753130ea752c70ff3d6b2f1462f
Covariance computation wants to do a triangular iteration but as a
single loop. Right now it iterates over a square and does nothing half
the time, which is inefficient and has bad worst-case threading
performance. This adds a utility that allows waste-free linear iteration
over a triangle.
Change-Id: I881d5683c65882f87dc2b5f8449a855d22ace755
This converts the bundle_adjustment_test to the parallel
version with multiple binaries, as is already in place for
the Bazel build. Additionally, since there is no longer a
need for it, this deletes bundle_adjustment_test.cc.
The test suite now runs on my 6 year old desktop in ~60 seconds!
Change-Id: Ib4a59f8749e823f697e6da5d303977d284303ae3
Previously, the thread ID was acquired and released on every iteration
of the for loop. The C++11 concurrent queue implementation is much
slower than TBB's version and consequently this was a huge bottleneck.
This introduces another ParallelFor API which takes the thread ID as a
parameter in the evaluation function. This allows us to acquire and
release the thread ID for each block of work which drastically improves
the performance.
This change brings us on par with OpenMP and TBB. See below for a
timing comparison. Note: in this example this CLs C++11 version is
faster to compute the residuals because TBB still must acquire the
thread ID on every iteration, which has some overhead.
Tested by building and running tests for no threading, OpenMP, TBB, and
C++11 threads. Also ran bazel tests.
./bin/bundle_adjuster --input=problem-744-543562-pre.txt --num_threads=8
C++11 @Head
Time (in seconds):
Residual only evaluation 7.819692 (5)
Jacobian & residual evaluation 11.606063 (6)
Linear solver 47.860195 (5)
Minimizer 70.877072
Total 90.806338
---------------------------------------------------
C++11 (This CL)
Time (in seconds):
Residual only evaluation 1.217500 (5)
Jacobian & residual evaluation 5.796112 (6)
Linear solver 44.080873 (5)
Minimizer 54.635524
Total 77.640072
---------------------------------------------------
OpenMP
Time (in seconds):
Residual only evaluation 0.797023 (5)
Jacobian & residual evaluation 5.633916 (6)
Linear solver 43.280020 (5)
Minimizer 53.199058
Total 76.250861
---------------------------------------------------
TBB
Time (in seconds):
Residual only evaluation 1.911095 (5)
Jacobian & residual evaluation 5.557807 (6)
Linear solver 44.074680 (5)
Minimizer 55.002688
Total 78.052687
---------------------------------------------------
No Threads
Time (in seconds):
Residual only evaluation 2.939212 (5)
Jacobian & residual evaluation 18.519874 (6)
Linear solver 74.017837 (5)
Minimizer 98.980080
Total 122.216391
Change-Id: I3af959b0771bbdfe8cad8c13896191d6ac903181
This change is provided on behalf of Steve Hsu.
Tested by compiling and inspecting the documentation.
Change-Id: Ib892bcc3ad76cba1bad133a1fd1d26468d0e6437
This change while algebraically equivalent was causing non-trivial
changes in the numbers being reported by ceres and in some cases
they got worse.
Also fix a small header include in parallel_for_test.cc which
was discovered as part of testing this patch.
Change-Id: I8c8d61538819f0b25af6059fa75ad068fc22f5e4
1. Solver::Options::num_threads now controls parallelism in Ceres
Solver. The user specified value of
Solver::Options::num_linear_solver_threads is ignored.
2. If the user specifies Solver::Options::num_linear_solver_threads
and it is different from Solver::Options::num_threads,
a warning is printed.
3. Solver::Summary:num_linear_solver_threads_given and
Solver::Summary::num_linear_solver_threads_used are also
deprecated and are always set to Solver::Summary::num_threads_given
and Solver::Summary::num_threads_used.
Change-Id: I20b9336d9336e400e6f0a15b63857c0c43eb271c
This improves the readability and simplifies the logic
for interfacing with ParallelFor. More importantly, it paves the
way for ParallelFor refactoring to improve its performance.
Change-Id: I13b05596228900ee00d71f2ccce1db338844b9ab