Also removed Solver::Summary::num_linear_solver_threads_given
and Solver::Summary::num_linear_solver_threads_used.
Change-Id: I559145ae2e7af597ea06ec03d386645a3a892e9f
- As OPENMP, TBB & CXX11_THREADS are mutually exclusive, and their
implementations are wholly contained within enclosing #ifdefs if the
requisite option is not enabled then the source file is 'empty'.
- This triggers 'xxx.cc has no symbol' warnings when compiling
Ceres statically, as such we now only include these files if they are
not going to be empty.
Change-Id: I5390251b44c5770bbdd4ccb1bc916dbe3b1814a9
The way the SystemTest fixture works is that it takes
a "FooProblem" object as a type, which contains a ceres::Problem
and a ceres::Solver::Options object.
The Options object also contains a linear_solver_ordering which
contains double* which refer to memory that is allocated when
a problem object is created.
So it is important that the lifetime of the ceres::Problem object
and the ceres::Solver::Options object be tied together. But we were
violating this by creating a FooProblem object on the stack, grabbing
its Options struct and passing it to the SystemTest fixture, which
would then create another instance of FooProblem, grab its Problem
object and copy the modified options struct into it.
In the case where a user provided ordering was being used,
this ordering would now be referring to memory allocated by the first
FooProblem object, which would cause Ceres's internal ApplyOrdering
function to fail.
The fix is ofcourse to Problem and Options object that are born
together.
Change-Id: I07c377a9d5fcabbb6c7ca8aa3460206ce045ffa9
1. Remove SolverConfig, this was a wrapper
around Solver::Options. As we experiment
with more Solver::Options, it became a hurdle.
2. Updated generate_bundle_adjustment_tests.py to use
Solver::Options directly.
3. Update system_test to use Solver::Options.
NOTE: generate_bundle_adjustment_tests.py changes are a bit
gunky, but I tried to minimize the changes in this CL
as I am going to introduce new test cases and that
is going to significantly change this file.
Change-Id: I34a2f51824b04ef368a5bbe54fbd7b281381909e
1. Replace CERES_DISALLOW_* with explicitly deleted constructors.
2. Replace use of CERES_ARRAY_SIZE and stack allocated arrays
with std::vector.
3. Move CERES_ALIGN_* macros into manual_constructor.h, which is
the one place they are used and will be deprecated along with that
file.
4. Introduce isnan,isnormal,isinf and isfinite for Jets.
5. Replace IsNormal,IsFinite,IsNaN and IsInfinite with corresponding
c++11 function calls.
Change-Id: I04f33a221aae77d247602150988b6d4aa4efeeab
Change the loop structure of IterativeRefiner to
unconditionally refine for max_num_iterations.
This is done for two reasons.
1. We expect to use this refinement for a small number of iterations
where the convergence test is useless.
2. Eliminating the convergence test means we can restructure the loop
and save on a sparse matrix-vector multiply, saving precious
compute.
Change-Id: I6347f453a5d19d234af2a2eb1bce811048963e06
1. Add Solver::Options::use_mixed_precision_solves,
and Solver::Options::max_num_refinement_iterations.
2. Make SparseCholesky::Create return a unique_ptr.
3. SparseCholesky::Create now takes LinearSolver::Options
as an argument.
4. IterativeRefiner's constructor does not require num_cols
as an argument.
5. SparseNormalCholeskySolver now uses a separate rhs vector.
This basic implementation results in a 10% reduction in solver time
and 30% reduction in linear solver memory usage.
Change-Id: I6830f32cae2febf082d2733262eb2c9f0482b0ea
1. Add SuiteSparse::CreateDenseVectorView
2. Replace calls to SuiteSparse::CreateDenseVector with
SuiteSparse::CreateDenseVectorView.
2. Replace NULL with nullptr in suitesparse.cc and
dynamic_sparse_normal_cholesky_solver.cc
Change-Id: I94355c1dc27789e5b987a7b2850e9db6176a0914
With the addition of C++11 support we can simplify the parallel for code by
removing the ifdef branching. Converts coordinate_descent_minimizer.cc to use
the thread_id ParallelFor API.
Tested by building with OpenMP, C++11 threads, TBB, and no threads. All tests
pass.
Also compared timing via the bundle adjuster.
./bin/bundle_adjuster --input=../problem-744-543562-pre.txt
With OpenMP num_threads=8
Head:
Time (in seconds):
Residual only evaluation 0.807753 (5)
Jacobian & residual evaluation 4.489404 (6)
Linear solver 41.826481 (5)
Minimizer 50.745857
Total 73.294424
CL:
Time (in seconds):
Residual only evaluation 0.970483 (5)
Jacobian & residual evaluation 4.647438 (6)
Linear solver 41.781892 (5)
Minimizer 50.848904
Total 73.089983
With OpenMP num_threads=1
HEAD:
Time (in seconds):
Residual only evaluation 2.990246 (5)
Jacobian & residual evaluation 14.132090 (6)
Linear solver 79.631951 (5)
Minimizer 100.281847
Total 122.946267
CL:
Time (in seconds):
Residual only evaluation 3.075178 (5)
Jacobian & residual evaluation 13.966451 (6)
Linear solver 77.005441 (5)
Minimizer 97.568712
Total 120.410454
Change-Id: I1857d7943073be7465b6c6476bf46ab11c5475a3
1. Add explicit tests for MatrixMatrixMultiplyNaive and
MatrixTransposeMatrixMultiplyNaive
2. Add tests that exercise a variety of matrix sizes for
MatrixVectorMultiply and MatrixTransposeVectorMultiply.
Change-Id: I0b25ec346b719f19b2067848f9d4bb64c9848750
Migrate all Option and Summary structs to use
inline member initialization syntax.
This reduces the amount of code, and collocates the
default values with the documentation for the corresponding
member variable.
Change-Id: I8e6b9ee3b31464699d678667f6166ace5fc137c9
- Previously we were requesting a step-size which minimised the
polynomial in: [previous.x, current.x * factor].
- This is incorrect c/f Nocedal & Wright p60, the bounds for the
minimising step-size should be: [current.x, current.x * factor].
- However, Nocedal & Wright's bounds are insufficient when the function
can return invalid values, which we support.
- In the case that f(current) is invalid, we need to contract the
step-size to lie within: [previous.x, current.x) given that we know
that previous.x is valid so a valid step must exist within that range.
Change-Id: I67b5aaa09cc6d54cf5f264e2cf894ddc2af3f3ad
1. Replace HashMap and HashSet with std::unordered_map and
std::unordered_set respectively.
2. Extract the pair hasher into a struct pair_hash.
3. Delete collections_port.h
4. Convert explicit iterator based loops to auto based
loops where sensible.
Change-Id: Ib88bcd13a7463d18435639d3b771abaa52080efb
- Removes CXX11 option, and all associate paraphernalia. Ceres now
requires a compiler with full >= C++11 support. In MSVC terms this
means >= 2013 Release 4.
- This deprecates the use of CERES_STD_UNORDERED_MAP and CERES_USE_CXX11
as they will now always be defined. They will be removed from the
source in a future CL.
- For clients with CMake >= 3.8 we propagate via the exported/installed
Ceres target the CXX version that was specified when Ceres was built.
For versions < 3.8 (but >= 3.5) we specify the CXX features currently
used in the Ceres public API.
Change-Id: I535b545b10156e4426659c270a4a0649e071df0e
Even though we added support for storing the upper and
lower triangular parts of symmetric matrices in
CompressedRowSparseMatrix. RightMultiply, LeftMultiply
and SquaredColumnNorm were not modified to account for this.
This CL changes their implementation and adds thorough
tests.
Also methods that cannot work correctly with symmetric
storage now CHECK and fail, because that indicates
programmer error.
Change-Id: I76288472c8bac98db7376a79bdb6259e346ef2b7
This adds a callback mechanism to for users to get notified just
before jacobian and residual evaluations. This will enable
aggressive caching and sharing of compute between cost functions.
Change-Id: I67993726920218edf71ab9ae70c34c204756c71a
1. When Solver::Options::update_state_every_iteration = true,
the StateUpdatingCallback only updates the state on a
successful iteration. This works fine but will not work
if the user provides an EvaluationCallback, because
every call to Evaluator::Evaluate will change the state
visible to the user. This means that on unsuccessful
iterations when the user's IterationCallback is called
the state visible to the user would be the last evaluation
which did not lead to an improved cost. This CL forces
the StateUpdatingCallback to unconditionally update
the user visible state.
2. When the minimizer terminates, we update the user visible
state only if the solution is usable, otherwise the user's
state will remain the same. To maintain this invariant
we will now cache the user's state and make sure we
use it to reset it upon return if the solver is not
successful.
Change-Id: Ic1f8fa6cb10d2753130ea752c70ff3d6b2f1462f
Covariance computation wants to do a triangular iteration but as a
single loop. Right now it iterates over a square and does nothing half
the time, which is inefficient and has bad worst-case threading
performance. This adds a utility that allows waste-free linear iteration
over a triangle.
Change-Id: I881d5683c65882f87dc2b5f8449a855d22ace755
This converts the bundle_adjustment_test to the parallel
version with multiple binaries, as is already in place for
the Bazel build. Additionally, since there is no longer a
need for it, this deletes bundle_adjustment_test.cc.
The test suite now runs on my 6 year old desktop in ~60 seconds!
Change-Id: Ib4a59f8749e823f697e6da5d303977d284303ae3
Previously, the thread ID was acquired and released on every iteration
of the for loop. The C++11 concurrent queue implementation is much
slower than TBB's version and consequently this was a huge bottleneck.
This introduces another ParallelFor API which takes the thread ID as a
parameter in the evaluation function. This allows us to acquire and
release the thread ID for each block of work which drastically improves
the performance.
This change brings us on par with OpenMP and TBB. See below for a
timing comparison. Note: in this example this CLs C++11 version is
faster to compute the residuals because TBB still must acquire the
thread ID on every iteration, which has some overhead.
Tested by building and running tests for no threading, OpenMP, TBB, and
C++11 threads. Also ran bazel tests.
./bin/bundle_adjuster --input=problem-744-543562-pre.txt --num_threads=8
C++11 @Head
Time (in seconds):
Residual only evaluation 7.819692 (5)
Jacobian & residual evaluation 11.606063 (6)
Linear solver 47.860195 (5)
Minimizer 70.877072
Total 90.806338
---------------------------------------------------
C++11 (This CL)
Time (in seconds):
Residual only evaluation 1.217500 (5)
Jacobian & residual evaluation 5.796112 (6)
Linear solver 44.080873 (5)
Minimizer 54.635524
Total 77.640072
---------------------------------------------------
OpenMP
Time (in seconds):
Residual only evaluation 0.797023 (5)
Jacobian & residual evaluation 5.633916 (6)
Linear solver 43.280020 (5)
Minimizer 53.199058
Total 76.250861
---------------------------------------------------
TBB
Time (in seconds):
Residual only evaluation 1.911095 (5)
Jacobian & residual evaluation 5.557807 (6)
Linear solver 44.074680 (5)
Minimizer 55.002688
Total 78.052687
---------------------------------------------------
No Threads
Time (in seconds):
Residual only evaluation 2.939212 (5)
Jacobian & residual evaluation 18.519874 (6)
Linear solver 74.017837 (5)
Minimizer 98.980080
Total 122.216391
Change-Id: I3af959b0771bbdfe8cad8c13896191d6ac903181
This change while algebraically equivalent was causing non-trivial
changes in the numbers being reported by ceres and in some cases
they got worse.
Also fix a small header include in parallel_for_test.cc which
was discovered as part of testing this patch.
Change-Id: I8c8d61538819f0b25af6059fa75ad068fc22f5e4
1. Solver::Options::num_threads now controls parallelism in Ceres
Solver. The user specified value of
Solver::Options::num_linear_solver_threads is ignored.
2. If the user specifies Solver::Options::num_linear_solver_threads
and it is different from Solver::Options::num_threads,
a warning is printed.
3. Solver::Summary:num_linear_solver_threads_given and
Solver::Summary::num_linear_solver_threads_used are also
deprecated and are always set to Solver::Summary::num_threads_given
and Solver::Summary::num_threads_used.
Change-Id: I20b9336d9336e400e6f0a15b63857c0c43eb271c
This improves the readability and simplifies the logic
for interfacing with ParallelFor. More importantly, it paves the
way for ParallelFor refactoring to improve its performance.
Change-Id: I13b05596228900ee00d71f2ccce1db338844b9ab
This reverts commit 6c835f81a7.
This CL introduced a subtle cancellation error which is causing
some application tests to fail, and is therefore being reverted.
Change-Id: Iaaa8f13a18f0a805b0e70a25ccd1e97efb21f330
The enum Eigen::UpLoType did not have a name in older versions
of Eigen, so a templated function using that enum type fails
to compile with earlier versions of Eigen.
This change replaces the enum in the template declaration with
an int.
Change-Id: Id128fd96b76818be347ee6ed5945c231936d9af8
Implements ParallelFor using the C++11 based ThreadPool. The C++11
parallel for is 50-70% faster than single threaded, and 20-30% slower
than TBB.
Tested by compiling with OpenMP, TBB, and C++11 Threading support and
ran the unit tests. Ran bazel as well.
Change-Id: I7fd6c9037ff9f200ce6999b5f39918995bb6b8ea