wip: adaptive-rate frame extraction from videos

This commit is contained in:
Harry Chen
2026-09-17 00:17:05 -04:00
parent bb9116439d
commit d0e3673d13
28 changed files with 1917 additions and 192 deletions
+10 -1
View File
@@ -65,6 +65,7 @@ set(SS_TOOL_LIBS "")
list(APPEND SS_TOOL_SOURCES
${SS_SRC}/app/FrameMask.cpp
${SS_SRC}/app/FrameLook.cpp
${SS_SRC}/app/FrameMotion.cpp
${SS_SRC}/app/Pano360.cpp
${SS_SRC}/app/AppPaths.cpp
${SS_SRC}/app/CrashLog.cpp)
@@ -309,7 +310,8 @@ if(SS_SEPARATE_TOOLS)
endif()
if(SS_BUILD_SAM)
set(_sam_src ${SS_SRC}/app/cli/sam_main.cpp ${SS_SRC}/app/FrameMask.cpp
${SS_SRC}/app/FrameLook.cpp ${SS_SRC}/app/Pano360.cpp)
${SS_SRC}/app/FrameLook.cpp ${SS_SRC}/app/FrameMotion.cpp
${SS_SRC}/app/Pano360.cpp)
set(_sam_lib ss_sam)
if(SS_ENABLE_PATENTED)
list(APPEND _sam_src ${SS_SRC}/app/cli/sam_extract.cpp
@@ -332,6 +334,13 @@ foreach(test_src ${SS_CORE_TESTS})
ss_configure_app(${test_name})
endforeach()
# The frame plan: no device, no GUI, and a wrong answer is silent.
add_executable(frame_motion_test
${SS_SRC}/app/tests/frame_motion_test.cpp
${SS_SRC}/app/FrameMotion.cpp
${SS_SRC}/app/Pano360.cpp)
ss_configure_app(frame_motion_test)
# The GUI files with no GUI in them: the stamp that decides whether a finished
# reconstruction is kept or built again, and the preset serializers. Named
# rather than globbed -- each such test names its own sources.
+67 -1
View File
@@ -462,11 +462,12 @@ video modes, 0 and 4 for timelapse -- enumerate them, do not hardcode).
Together they are a YouTube-style **EAC 3x2 cubemap**: the first track is the
top row (LEFT, FRONT, RIGHT), the second the bottom (DOWN rot270, BACK rot90,
UP rot270). Faces are square and as tall as a track. Each track carries three
of them plus **two 32 px overlap strips**, inserted at the centre lines of its
of them plus **two overlap strips**, inserted at the centre lines of its
two side faces, which is where the two lenses meet:
| mode | track | face | strips | canvas |
|---|---|---|---|---|
| 8K | 5952x1920 | 1920 | 2 x 96 | 5760x3840 |
| 5.6K | 4096x1344 | 1344 | 2 x 32 | 4032x2688 |
| 3K | 2272x736 | 736 | 2 x 32 | 2208x1472 |
@@ -553,6 +554,71 @@ high`) -- about 4.4 px per degree, against 16.7 for a 1504 px face.
- **The operator.** Whoever is holding it is in the downward and rearward views
of every frame of most captures, and wants masking out.
## Frames out of a video
One frame every `skip` source frames, the sharpest of a window of `keep`
around each. Both decode paths choose the same frames (`app/FrameExtract.h`),
and the GUI's rate is **per input**: a capture shot as several clips is rarely
shot at one pace, so `PrepInput::fps` overrides the job's for that video.
### Adaptive spacing
`--adaptive` (the GUI: "Adapt the rate to the motion") spaces the kept frames
by how much the **view** changed instead of by how much time passed. The rate
asked for becomes the average; the realized rate stays within `--adaptive-range`
either side of it, and the frame count comes out the same.
`app/FrameMotion.h` measures it. A grid of about 500 points is tracked between
small grey frames by pyramidal Lucas-Kanade, one global model is fitted to the
flow by RANSAC, and two numbers come out:
- **coverage** -- the share of the frame's content that left it, as the area of
the frame carried over by the fitted model and cut back to the frame. Pan and
zoom cost; roll costs only its corners, which is right, because a rolled
frame still sees what it saw.
- **parallax** -- the flow the global model could not explain, as the 75th
percentile of the residual. This is what a translation past something close
produces and what a translation towards something far does not.
`cost = coverage + 2 * parallax`, and the plan spaces frames at equal
cumulative cost. The weight is the only hand-set number: a tenth of the frame
of unexplained disparity is a harder match than a tenth of the frame of pan,
and carries the triangulation the pan does not.
The global model depends on what the frames are pictures of:
| capture | model | coverage |
|---|---|---|
| ordinary video | 2D affine, in the image | the frame carried over |
| `.360` (EAC) | a rotation of the SPHERE, through the packing's own mapping | 0 |
| dual fisheye (`.insv`, `.osv`) | a rotation of the sphere, equidistant lens | 0 |
A camera that sees every direction keeps every direction when it turns, so a
360 capture spinning on the spot scores nothing and is given frames at the
slowest rate the bounds allow. That is the whole point of fitting the rotation
in 3D rather than fitting a homography per track: on a sphere, turning is free
and only moving is not.
Two details that were measured rather than chosen:
- A sphere frame is analyzed at **four times the pixels** of a flat one. It
spans three times the angle, and at flat resolution the tracking floor is
itself a degree wide and drowns the parallax.
- The step is found by **bisection** on the wanted count, not as
`total / count`. A burst of motion swallows several steps' budget within one
sample and can spend only one frame of it, which left plans a third short.
It costs one extra decode pass over the first track of each video: measured at
+21% on a dual-fisheye `.osv` (whose main pass already decodes two tracks) and
up to +100% on a 1080p clip, with the tracking itself overlapped with the
decode. Nothing is buffered: a video's worth of pictures does not fit, and a
plan cannot be made until the whole cost curve is known.
Without the built-in decoder the same plan is made from the candidate frames
ffmpeg already extracts, at `fps x max(window, range)` instead of
`fps x window` so there are enough of them for the fastest rate it may ask for
(`gui/FrameSelect.h`).
## Preprocessing tools
`reference/scripts/` holds standalone Python utilities that produce these
+172 -26
View File
@@ -6,6 +6,7 @@
#include "app/FrameExtract.h"
#include "app/FrameMotion.h"
#include "app/WriterPool.h"
#include "i18n/catalog/Data.h"
#include "i18n/catalog/Log.h"
@@ -17,6 +18,7 @@
#include <algorithm>
#include <atomic>
#include <cmath>
#include <condition_variable>
#include <cstdio>
#include <deque>
@@ -48,22 +50,150 @@ using app::WriterPool;
// because the paired path has to choose the same frames as the single-track
// one and as the ffmpeg fallback.
struct SelectClock {
// A fixed rate is arithmetic on the index; `ends`, from the motion pass,
// is the frame each window closes at. Both are walked forward only.
int keep = 0, skip = 1;
int64_t count = 0;
bool candidate(int64_t i) const {
std::vector<int64_t> ends;
size_t next = 0;
bool candidate(int64_t i) {
if (!ends.empty()) {
while (next < ends.size() && ends[next] < i) ++next;
return next < ends.size() && i > ends[next] - std::max(keep, 1) &&
i <= ends[next];
}
return keep == 0 ? (i % skip == 0) : (((i + keep) % skip) < keep);
}
void skipped() { ++count; }
// The Python increments its counter before the modulo test when a blur
// window is in use, and after it when it is not; both are reproduced here.
bool step() {
if (keep != 0) ++count;
const bool write = (count % skip) == 0;
if (keep == 0) ++count;
return write;
bool step(int64_t i) {
if (!ends.empty()) return next < ends.size() && i == ends[next];
return keep == 0 || ((i + 1) % skip) == 0;
}
};
// ---------------------------------------------------------------------------
// Adaptive selection
// ---------------------------------------------------------------------------
// Pass one of an adaptive run: ONE track, reduced on the GPU so a frame costs
// a few tens of kilobytes over the bus. Decoding twice is cheaper than holding
// a video's worth of pictures until the plan is known.
bool measure_motion(const FrameExtractJob& o, const FrameExtractSinks& sinks,
int track, int tracks, std::vector<int64_t>& plan,
FrameExtractStats& t, std::string& error) {
const double t_start = nn::now_ms();
video::VideoPipeline pipe;
if (!pipe.open(o.input, track, 1, error)) return false;
const video::TrackInfo& info = pipe.track();
MotionOptions mo;
// A capture that sees the whole sphere: the 360 packing, or the two square
// tracks every dual-fisheye camera writes.
if (o.eac.valid()) mo.view = MotionView::Eac360;
else if (tracks >= 2 && info.width == info.height)
mo.view = MotionView::Fisheye;
mo.eac = o.eac;
mo.out_fov = mo.view == MotionView::Fisheye ? mo.circle_fov
: motion_out_fov(o.views);
motion_frame_size(mo.view, info.width, info.height, mo.width, mo.height);
MotionTracker tracker(mo);
const double fps = info.fps > 1.0 ? info.fps : 30.0;
// 15 samples a second is enough to follow a walk, but never coarser than a
// third of the spacing being planned: a gap can only land on a sample.
const int stride = std::max(
1, std::min((int)std::lround(fps / 15.0), std::max(1, o.skip / 3)));
// Tracking is CPU and decoding is GPU, so they run at the same time: the
// decoder fills a short queue and one consumer thread owns the tracker.
struct Sample {
std::vector<uint8_t> gray;
int64_t index = 0;
};
std::deque<Sample> queue;
std::mutex mu;
std::condition_variable room, work;
bool feeding = true;
std::thread consumer([&] {
for (;;) {
Sample s;
{
std::unique_lock<std::mutex> lk(mu);
work.wait(lk, [&] { return !queue.empty() || !feeding; });
if (queue.empty()) return;
s = std::move(queue.front());
queue.pop_front();
}
room.notify_one();
tracker.track(s.gray.data(), s.index);
}
});
auto stop = [&]() {
{
std::lock_guard<std::mutex> lk(mu);
feeding = false;
}
work.notify_all();
consumer.join();
};
std::vector<uint8_t> gray;
int64_t last = -1;
for (;;) {
if (sinks.cancel && sinks.cancel->load()) {
error = "cancelled";
stop();
return false;
}
video::FrameHandle h;
if (!pipe.next(h, error)) {
if (!error.empty()) {
stop();
return false;
}
break;
}
if (h.index % stride == 0) {
if (!pipe.toGray(h, mo.width, mo.height, gray, error)) {
pipe.release(h);
stop();
return false;
}
std::unique_lock<std::mutex> lk(mu);
room.wait(lk, [&] { return queue.size() < 3; });
queue.push_back({gray, h.index});
lk.unlock();
work.notify_one();
++t.analyzed;
}
last = h.index;
pipe.release(h);
}
stop();
tracker.finish();
const int64_t frames = info.frame_count > 0 ? info.frame_count : last + 1;
plan = plan_by_motion(tracker.costs(), tracker.ends(), frames, o.skip,
std::max(o.keep, 1), o.adaptive_range, o.max_frames);
t.plan = nn::now_ms() - t_start;
if (plan.empty()) {
error = "the capture is too short to space frames by motion";
return false;
}
int64_t tightest = frames, widest = 0;
for (size_t i = 1; i < plan.size(); i++) {
const int64_t gap = plan[i] - plan[i - 1];
tightest = std::min(tightest, gap);
widest = std::max(widest, gap);
}
auto rate = [&](int64_t gap) {
return std::round((gap > 0 ? fps / (double)gap : fps) * 100.0) / 100.0;
};
namespace lmsg = spirula::i18n::msg::log;
log_line(sinks,
spirula::i18n::format(lmsg::motion_plan, {(long long)plan.size(),
rate(widest), rate(tightest)}));
return true;
}
// ---------------------------------------------------------------------------
// Extraction
// ---------------------------------------------------------------------------
@@ -71,8 +201,8 @@ struct SelectClock {
bool extract_track(const FrameExtractJob& o, const FrameExtractSinks& sinks,
int track, const fs::path& image_dir,
const fs::path& mask_dir, sam::Masker* masker,
WriterPool& pool, FrameExtractStats& t,
std::string& error) {
WriterPool& pool, const std::vector<int64_t>& plan,
FrameExtractStats& t, std::string& error) {
video::VideoPipeline pipe;
// The blur window holds `keep` decoded pictures at once.
if (!pipe.open(o.input, track, std::max(o.keep, 1), error)) return false;
@@ -90,7 +220,7 @@ bool extract_track(const FrameExtractJob& o, const FrameExtractSinks& sinks,
};
std::deque<Buffered> window;
const int keep = o.keep;
SelectClock clock{o.keep, o.skip, 0};
SelectClock clock{o.keep, o.skip, plan, 0};
int written = 0;
bool measured_pending = false;
@@ -192,7 +322,6 @@ bool extract_track(const FrameExtractJob& o, const FrameExtractSinks& sinks,
// else is decoded (inter prediction needs it) but never touched again.
if (!clock.candidate(i)) {
pipe.release(h);
clock.skipped();
continue;
}
@@ -209,7 +338,7 @@ bool extract_track(const FrameExtractJob& o, const FrameExtractSinks& sinks,
window.pop_front();
}
if (clock.step()) flush_window(true);
if (clock.step(i)) flush_window(true);
}
flush_window(false);
return true;
@@ -223,7 +352,8 @@ using LockstepSink = std::function<bool(std::vector<nn::Image>& imgs, int64_t in
bool extract_lockstep(const FrameExtractJob& o, const FrameExtractSinks& sinks,
const std::vector<int>& tracks, const video::ConvertOpts& conv,
FrameExtractStats& t, std::string& error, const LockstepSink& on_frame) {
const std::vector<int64_t>& plan, FrameExtractStats& t,
std::string& error, const LockstepSink& on_frame) {
const size_t n = tracks.size();
std::vector<std::unique_ptr<video::VideoPipeline>> pipe;
for (size_t k = 0; k < n; k++) {
@@ -237,7 +367,7 @@ bool extract_lockstep(const FrameExtractJob& o, const FrameExtractSinks& sinks,
};
std::deque<Buffered> window;
const int keep = o.keep;
SelectClock clock{o.keep, o.skip, 0};
SelectClock clock{o.keep, o.skip, plan, 0};
int written = 0;
bool measured_pending = false;
@@ -317,7 +447,6 @@ bool extract_lockstep(const FrameExtractJob& o, const FrameExtractSinks& sinks,
b.index = b.h[0].index;
if (!clock.candidate(b.index)) {
release(b);
clock.skipped();
continue;
}
if (keep != 0) {
@@ -332,7 +461,7 @@ bool extract_lockstep(const FrameExtractJob& o, const FrameExtractSinks& sinks,
release(window.front());
window.pop_front();
}
if (clock.step()) flush_window(true);
if (clock.step(b.index)) flush_window(true);
}
flush_window(false);
return true;
@@ -342,7 +471,8 @@ bool extract_lockstep(const FrameExtractJob& o, const FrameExtractSinks& sinks,
// into every view.
bool extract_pair(const FrameExtractJob& o, const FrameExtractSinks& sinks,
const fs::path& image_dir, WriterPool& pool,
FrameExtractStats& t, std::string& error) {
const std::vector<int64_t>& plan, FrameExtractStats& t,
std::string& error) {
{
std::vector<std::pair<int, int>> sizes = video_track_sizes(o.input, error);
for (size_t k = 0; k < 2; k++) {
@@ -396,14 +526,16 @@ bool extract_pair(const FrameExtractJob& o, const FrameExtractSinks& sinks,
}
return true;
};
return extract_lockstep(o, sinks, {0, 1}, video::ConvertOpts{}, t, error, on_frame);
return extract_lockstep(o, sinks, {0, 1}, video::ConvertOpts{}, plan, t, error,
on_frame);
}
// Every track of a multi-lens file at the same instants, one folder per
// track: the rig the frames of one stem form is then a rig in fact.
bool extract_synced(const FrameExtractJob& o, const FrameExtractSinks& sinks,
const std::vector<int>& tracks, const fs::path& base, WriterPool& pool,
FrameExtractStats& t, std::string& error) {
const std::vector<int64_t>& plan, FrameExtractStats& t,
std::string& error) {
std::vector<fs::path> dirs;
for (size_t k = 0; k < tracks.size(); k++) {
dirs.push_back(base / ("cam" + std::to_string(k)));
@@ -431,7 +563,7 @@ bool extract_synced(const FrameExtractJob& o, const FrameExtractSinks& sinks,
}
return true;
};
return extract_lockstep(o, sinks, tracks, conv, t, error, on_frame);
return extract_lockstep(o, sinks, tracks, conv, plan, t, error, on_frame);
}
} // namespace
@@ -565,11 +697,19 @@ bool extract_frames(const FrameExtractJob& job_in, const FrameExtractSinks& sink
const double t_start = nn::now_ms();
bool ok = true;
const bool synced = !pano && job.sync_tracks && tracks.size() > 1;
// One plan for the whole file, measured on its first track: the tracks of
// a rig see the same motion, and the ones that do not are still one camera
// moving through one scene.
std::vector<int64_t> plan;
if (job.adaptive &&
!measure_motion(job, sinks, tracks[0], (int)tracks.size(), plan, stats,
error))
return false;
if (pano) {
stats.tracks = 2;
ok = extract_pair(job, sinks, base, pool, stats, error);
ok = extract_pair(job, sinks, base, pool, plan, stats, error);
} else if (synced) {
ok = extract_synced(job, sinks, tracks, base, pool, stats, error);
ok = extract_synced(job, sinks, tracks, base, pool, plan, stats, error);
}
for (size_t ti = 0; ti < tracks.size() && ok && !pano && !synced; ++ti) {
// A multi-track file (an Insta360 .insv carries two fisheye streams)
@@ -582,7 +722,7 @@ bool extract_frames(const FrameExtractJob& job_in, const FrameExtractSinks& sink
log_line(sinks, "track " + std::to_string(tracks[ti]) + " -> " +
image_dir.string());
ok = extract_track(job, sinks, tracks[ti], image_dir, mask_dir, masker.get(),
pool, stats, error);
pool, plan, stats, error);
}
const double t_drain = nn::now_ms();
@@ -792,10 +932,16 @@ std::string format_extract_stats(const FrameExtractStats& s,
rows.push_back({dmsg::xs_written.get(),
format(dmsg::xs_written_to,
{(long long)s.written, out_dir}), false});
if (s.analyzed > 0)
rows.push_back({dmsg::xs_analyzed.get(),
std::to_string((long long)s.analyzed), false});
rows.push_back({dmsg::xs_decode.get(),
format(dmsg::xs_ms_fps,
{ms(s.decode), rate((double)s.decoded, s.decode)}),
true});
if (s.plan > 0)
rows.push_back({dmsg::xs_motion.get(),
format(dmsg::xs_ms, {ms(s.plan)}), false});
rows.push_back({dmsg::xs_sharpness.get(),
format(dmsg::xs_ms, {ms(s.sharpness)}), false});
rows.push_back({dmsg::xs_convert.get(),
+7 -1
View File
@@ -46,6 +46,11 @@ struct FrameExtractJob : FrameLook {
// instants (one sharpness window over all of them), so the frames of one
// stem are a rig. Off picks each track's sharpest frame on its own.
bool sync_tracks = false;
// Space the kept frames by view change instead of by time
// (app/FrameMotion.h): `skip` then sets the average and the rate stays
// within `adaptive_range` of it. Costs one extra pass over the first track.
bool adaptive = false;
float adaptive_range = 4.0f;
// Masking. Empty model = no masks.
sam::MaskOptions mask;
@@ -55,7 +60,8 @@ struct FrameExtractJob : FrameLook {
struct FrameExtractStats {
double decode = 0, sharpness = 0, convert = 0, mask = 0, submit = 0;
double drain = 0, total = 0, encode_cpu = 0;
int64_t decoded = 0, measured = 0, written = 0;
double plan = 0; // the adaptive pass, decode included
int64_t decoded = 0, measured = 0, written = 0, analyzed = 0;
int tracks = 1;
int encoder_threads = 0;
int write_failures = 0;
+706
View File
@@ -0,0 +1,706 @@
// FrameMotion.cpp -- see app/FrameMotion.h.
#include "app/FrameMotion.h"
#include <algorithm>
#include <cmath>
#include <cstring>
#include <thread>
namespace app {
namespace {
constexpr double kPi = 3.14159265358979323846;
// Residual flow weighs this much against content leaving the view. A tenth of
// the frame of unexplained disparity is a harder match than a tenth of the
// frame of pan, and carries the triangulation the pan does not.
constexpr float kParallaxWeight = 2.0f;
constexpr int kWin = 5; // 11x11 tracking window
constexpr int kWinArea = (2 * kWin + 1) * (2 * kWin + 1);
constexpr int kRansac = 64;
// ---------------------------------------------------------------------------
// Grey pyramids and Lucas-Kanade
// ---------------------------------------------------------------------------
struct Pyramid {
std::vector<std::vector<uint8_t>> level;
std::vector<int> w, h;
};
void build_pyramid(const uint8_t* src, int w, int h, Pyramid& p) {
p.w.assign(1, w);
p.h.assign(1, h);
p.level.resize(1);
p.level[0].assign(src, src + (size_t)w * h);
while (p.level.size() < 5 && p.w.back() >= 64 && p.h.back() >= 64) {
const int pw = p.w.back();
const int nw = pw / 2, nh = p.h.back() / 2;
p.level.resize(p.level.size() + 1);
const std::vector<uint8_t>& up = p.level[p.level.size() - 2];
std::vector<uint8_t>& d = p.level.back();
d.resize((size_t)nw * nh);
for (int y = 0; y < nh; y++)
for (int x = 0; x < nw; x++) {
const uint8_t* q = up.data() + (size_t)(2 * y) * pw + 2 * x;
d[(size_t)y * nw + x] =
(uint8_t)((q[0] + q[1] + q[pw] + q[pw + 1] + 2) / 4);
}
p.w.push_back(nw);
p.h.push_back(nh);
}
p.level.resize(p.w.size());
}
float sample(const uint8_t* im, int w, int h, float x, float y) {
x = std::min(std::max(x, 0.0f), (float)w - 1.001f);
y = std::min(std::max(y, 0.0f), (float)h - 1.001f);
const int x0 = (int)x, y0 = (int)y;
const float ax = x - x0, ay = y - y0;
const uint8_t* p = im + (size_t)y0 * w + x0;
const float t = p[0] + (p[1] - p[0]) * ax;
const float b = p[w] + (p[w + 1] - p[w]) * ax;
return t + (b - t) * ay;
}
// The displacement of one point from `a` to `b`, coarse to fine. Inverse
// compositional: the normal matrix comes from `a` and is built once per level.
bool track_point(const Pyramid& a, const Pyramid& b, float x, float y,
float& dx, float& dy) {
float gx[kWinArea], gy[kWinArea], ref[kWinArea];
float d0 = 0, d1 = 0, err = 0;
bool fitted = false;
const int top = (int)a.level.size() - 1;
for (int L = top; L >= 0; --L) {
if (L != top) {
d0 *= 2;
d1 *= 2;
}
const int w = a.w[(size_t)L], h = a.h[(size_t)L];
const float sc = 1.0f / (float)(1 << L);
const float px = x * sc, py = y * sc;
if (px < kWin + 2 || py < kWin + 2 || px > w - kWin - 3 ||
py > h - kWin - 3)
continue;
const uint8_t* ia = a.level[(size_t)L].data();
const uint8_t* ib = b.level[(size_t)L].data();
float gxx = 0, gxy = 0, gyy = 0;
int k = 0;
for (int j = -kWin; j <= kWin; j++)
for (int i = -kWin; i <= kWin; i++, k++) {
const float sx = px + i, sy = py + j;
gx[k] = 0.5f * (sample(ia, w, h, sx + 1, sy) -
sample(ia, w, h, sx - 1, sy));
gy[k] = 0.5f * (sample(ia, w, h, sx, sy + 1) -
sample(ia, w, h, sx, sy - 1));
ref[k] = sample(ia, w, h, sx, sy);
gxx += gx[k] * gx[k];
gxy += gx[k] * gy[k];
gyy += gy[k] * gy[k];
}
const float det = gxx * gyy - gxy * gxy;
const float lo =
0.5f * (gxx + gyy -
std::sqrt(std::max(0.0f, (gxx - gyy) * (gxx - gyy) +
4.0f * gxy * gxy)));
if (det <= 1e-6f || lo < 4.0f * kWinArea) continue;
for (int it = 0; it < 8; it++) {
float bx = 0, by = 0;
err = 0;
k = 0;
for (int j = -kWin; j <= kWin; j++)
for (int i = -kWin; i <= kWin; i++, k++) {
const float di =
ref[k] - sample(ib, w, h, px + i + d0, py + j + d1);
bx += gx[k] * di;
by += gy[k] * di;
err += std::fabs(di);
}
const float n0 = (gyy * bx - gxy * by) / det;
const float n1 = (gxx * by - gxy * bx) / det;
d0 += n0;
d1 += n1;
if (n0 * n0 + n1 * n1 < 0.01f) break;
}
fitted = true;
}
if (!fitted || err / kWinArea > 22.0f) return false;
dx = d0;
dy = d1;
return true;
}
// ---------------------------------------------------------------------------
// Global models
// ---------------------------------------------------------------------------
// b = A * [a 1], by least squares. False when the points are collinear.
bool solve_affine(const std::vector<float>& ax, const std::vector<float>& ay,
const std::vector<float>& bx, const std::vector<float>& by,
const int* pick, int n, float A[6]) {
double m[3][3] = {{0, 0, 0}, {0, 0, 0}, {0, 0, 0}};
double ru[3] = {0, 0, 0}, rv[3] = {0, 0, 0};
for (int k = 0; k < n; k++) {
const size_t i = (size_t)(pick ? pick[k] : k);
const double x = ax[i], y = ay[i], u = bx[i], v = by[i];
const double t[3] = {x, y, 1.0};
for (int r = 0; r < 3; r++) {
for (int c = 0; c < 3; c++) m[r][c] += t[r] * t[c];
ru[r] += u * t[r];
rv[r] += v * t[r];
}
}
const double det =
m[0][0] * (m[1][1] * m[2][2] - m[1][2] * m[2][1]) -
m[0][1] * (m[1][0] * m[2][2] - m[1][2] * m[2][0]) +
m[0][2] * (m[1][0] * m[2][1] - m[1][1] * m[2][0]);
if (std::fabs(det) < 1e-9) return false;
double inv[3][3];
inv[0][0] = (m[1][1] * m[2][2] - m[1][2] * m[2][1]) / det;
inv[0][1] = (m[0][2] * m[2][1] - m[0][1] * m[2][2]) / det;
inv[0][2] = (m[0][1] * m[1][2] - m[0][2] * m[1][1]) / det;
inv[1][0] = (m[1][2] * m[2][0] - m[1][0] * m[2][2]) / det;
inv[1][1] = (m[0][0] * m[2][2] - m[0][2] * m[2][0]) / det;
inv[1][2] = (m[0][2] * m[1][0] - m[0][0] * m[1][2]) / det;
inv[2][0] = (m[1][0] * m[2][1] - m[1][1] * m[2][0]) / det;
inv[2][1] = (m[0][1] * m[2][0] - m[0][0] * m[2][1]) / det;
inv[2][2] = (m[0][0] * m[1][1] - m[0][1] * m[1][0]) / det;
for (int r = 0; r < 3; r++) {
A[r] = (float)(inv[r][0] * ru[0] + inv[r][1] * ru[1] + inv[r][2] * ru[2]);
A[3 + r] = (float)(inv[r][0] * rv[0] + inv[r][1] * rv[1] + inv[r][2] * rv[2]);
}
return true;
}
void apply_affine(const float A[6], float x, float y, float& u, float& v) {
u = A[0] * x + A[1] * y + A[2];
v = A[3] * x + A[4] * y + A[5];
}
// A convex polygon of at most eight corners: the frame, or the frame carried
// over by the fitted model and cut back to it.
struct Poly {
float x[12], y[12];
int n = 0;
};
float poly_area(const Poly& p) {
double a = 0;
for (int i = 0, j = p.n - 1; i < p.n; j = i++)
a += (double)p.x[j] * p.y[i] - (double)p.x[i] * p.y[j];
return (float)std::fabs(a) * 0.5f;
}
Poly clip_to_frame(const Poly& in, float w, float h) {
auto side = [&](int e, float x, float y) -> float {
if (e == 0) return x;
if (e == 1) return w - x;
if (e == 2) return y;
return h - y;
};
Poly a = in, b;
for (int e = 0; e < 4 && a.n > 0; e++) {
b.n = 0;
for (int i = 0, j = a.n - 1; i < a.n; j = i++) {
const float dj = side(e, a.x[j], a.y[j]);
const float di = side(e, a.x[i], a.y[i]);
if ((di >= 0) != (dj >= 0) && b.n < 12) {
const float t = dj / (dj - di);
b.x[b.n] = a.x[j] + t * (a.x[i] - a.x[j]);
b.y[b.n] = a.y[j] + t * (a.y[i] - a.y[j]);
b.n++;
}
if (di >= 0 && b.n < 12) {
b.x[b.n] = a.x[i];
b.y[b.n] = a.y[i];
b.n++;
}
}
a = b;
}
return a;
}
struct Vec3 {
float x = 0, y = 0, z = 0;
};
Vec3 normalize(Vec3 v) {
const float n = std::sqrt(v.x * v.x + v.y * v.y + v.z * v.z);
if (n <= 0) return {0, 0, 1};
return {v.x / n, v.y / n, v.z / n};
}
Vec3 mul(const float R[9], const Vec3& v) {
return {R[0] * v.x + R[1] * v.y + R[2] * v.z,
R[3] * v.x + R[4] * v.y + R[5] * v.z,
R[6] * v.x + R[7] * v.y + R[8] * v.z};
}
float angle_between(const Vec3& a, const Vec3& b) {
const float dx = a.x - b.x, dy = a.y - b.y, dz = a.z - b.z;
const float c = std::sqrt(dx * dx + dy * dy + dz * dz) * 0.5f;
return 2.0f * std::asin(std::min(1.0f, c));
}
// The rotation two correspondences fix exactly: each pair spans an orthonormal
// frame, and one frame carried onto the other is the rotation.
bool rotation_from_two(const Vec3& a0, const Vec3& a1, const Vec3& b0,
const Vec3& b1, float R[9]) {
auto frame = [](const Vec3& u, const Vec3& v, float F[9]) {
const Vec3 e0 = normalize(u);
const float d = e0.x * v.x + e0.y * v.y + e0.z * v.z;
const Vec3 t{v.x - d * e0.x, v.y - d * e0.y, v.z - d * e0.z};
if (t.x * t.x + t.y * t.y + t.z * t.z < 1e-6f) return false;
const Vec3 e1 = normalize(t);
const Vec3 e2{e0.y * e1.z - e0.z * e1.y, e0.z * e1.x - e0.x * e1.z,
e0.x * e1.y - e0.y * e1.x};
F[0] = e0.x; F[1] = e1.x; F[2] = e2.x;
F[3] = e0.y; F[4] = e1.y; F[5] = e2.y;
F[6] = e0.z; F[7] = e1.z; F[8] = e2.z;
return true;
};
float Fa[9], Fb[9];
if (!frame(a0, a1, Fa) || !frame(b0, b1, Fb)) return false;
for (int r = 0; r < 3; r++)
for (int c = 0; c < 3; c++) {
float s = 0;
for (int k = 0; k < 3; k++) s += Fb[3 * r + k] * Fa[3 * c + k];
R[3 * r + c] = s;
}
return true;
}
// Horn's quaternion fit over the inliers, its largest eigenvector reached by
// shifted power iteration -- the matrix is 4x4, so an eigen solver would be
// more code than the thirty multiplies this is.
void rotation_from_many(const std::vector<Vec3>& a, const std::vector<Vec3>& b,
const std::vector<int>& pick, float R[9]) {
double S[3][3] = {{0, 0, 0}, {0, 0, 0}, {0, 0, 0}};
for (int i : pick) {
const float av[3] = {a[(size_t)i].x, a[(size_t)i].y, a[(size_t)i].z};
const float bv[3] = {b[(size_t)i].x, b[(size_t)i].y, b[(size_t)i].z};
for (int r = 0; r < 3; r++)
for (int c = 0; c < 3; c++) S[r][c] += (double)av[r] * bv[c];
}
const double N[4][4] = {
{S[0][0] + S[1][1] + S[2][2], S[1][2] - S[2][1], S[2][0] - S[0][2],
S[0][1] - S[1][0]},
{S[1][2] - S[2][1], S[0][0] - S[1][1] - S[2][2], S[0][1] + S[1][0],
S[2][0] + S[0][2]},
{S[2][0] - S[0][2], S[0][1] + S[1][0], -S[0][0] + S[1][1] - S[2][2],
S[1][2] + S[2][1]},
{S[0][1] - S[1][0], S[2][0] + S[0][2], S[1][2] + S[2][1],
-S[0][0] - S[1][1] + S[2][2]}};
double shift = 0;
for (int r = 0; r < 4; r++) {
double row = 0;
for (int c = 0; c < 4; c++) row += std::fabs(N[r][c]);
shift = std::max(shift, row);
}
double q[4] = {1, 0, 0, 0};
for (int it = 0; it < 40; it++) {
double n[4];
for (int r = 0; r < 4; r++) {
n[r] = shift * q[r];
for (int c = 0; c < 4; c++) n[r] += N[r][c] * q[c];
}
const double len =
std::sqrt(n[0] * n[0] + n[1] * n[1] + n[2] * n[2] + n[3] * n[3]);
if (len < 1e-12) break;
for (int r = 0; r < 4; r++) q[r] = n[r] / len;
}
const double w = q[0], x = q[1], y = q[2], z = q[3];
R[0] = (float)(1 - 2 * (y * y + z * z));
R[1] = (float)(2 * (x * y - w * z));
R[2] = (float)(2 * (x * z + w * y));
R[3] = (float)(2 * (x * y + w * z));
R[4] = (float)(1 - 2 * (x * x + z * z));
R[5] = (float)(2 * (y * z - w * x));
R[6] = (float)(2 * (x * z - w * y));
R[7] = (float)(2 * (y * z + w * x));
R[8] = (float)(1 - 2 * (x * x + y * y));
}
uint32_t next_random(uint32_t& s) {
s ^= s << 13;
s ^= s >> 17;
s ^= s << 5;
return s;
}
float percentile(std::vector<float>& v, float p) {
if (v.empty()) return 0.0f;
const size_t k = (size_t)std::max(
0.0, std::min((double)v.size() - 1, p * (double)(v.size() - 1)));
std::nth_element(v.begin(), v.begin() + (ptrdiff_t)k, v.end());
return v[k];
}
} // namespace
// ---------------------------------------------------------------------------
// The tracker
// ---------------------------------------------------------------------------
struct MotionTracker::Impl {
MotionOptions o;
float diag = 1.0f;
std::vector<float> gx, gy; // the grid tracked, in frame pixels
Pyramid prev, cur;
bool primed = false;
std::vector<float> cost;
std::vector<int64_t> end;
int weak = 0;
bool done = false;
struct Pair {
std::vector<float> ax, ay, bx, by;
};
bool sphere() const { return o.view != MotionView::Planar; }
float rad_per_pixel() const;
bool to_direction(float x, float y, Vec3& d) const;
float step_cost(const Pair& p) const;
float sphere_cost(const Pair& p) const;
float planar_cost(const Pair& p) const;
};
// How far apart two neighbouring grey pixels point, which is the floor every
// angular threshold here has to clear.
float MotionTracker::Impl::rad_per_pixel() const {
if (o.view == MotionView::Eac360) {
const float face = (float)o.width * (float)o.eac.face /
std::max(1.0f, (float)o.eac.track_w);
return (float)(kPi * 0.5) / std::max(face, 1.0f);
}
return 0.5f * o.circle_fov / std::max(o.circle_r, 1.0f);
}
bool MotionTracker::Impl::to_direction(float x, float y, Vec3& d) const {
if (o.view == MotionView::Eac360) {
// The grey frame is the whole track scaled down, strips and all, so
// the layout's own coordinates are what the mapping wants back.
const float sx = x * (float)o.eac.track_w / (float)o.width;
const float sy = y * (float)o.eac.track_h / (float)o.height;
float v[3];
if (!eac360_direction(o.eac, 0, sx, sy, v)) return false;
d = {v[0], v[1], v[2]};
return true;
}
const float rx = x - o.circle_x, ry = y - o.circle_y;
const float r = std::sqrt(rx * rx + ry * ry) / std::max(o.circle_r, 1.0f);
if (r > 1.0f) return false;
const float th = r * 0.5f * o.circle_fov, ph = std::atan2(ry, rx);
d = {std::sin(th) * std::cos(ph), std::sin(th) * std::sin(ph), std::cos(th)};
return true;
}
// Turning a camera that already sees every direction shows nothing new, so
// what the rotation leaves behind is the whole of the cost.
float MotionTracker::Impl::sphere_cost(const Pair& p) const {
const int n = (int)p.ax.size();
std::vector<Vec3> a, b;
a.reserve((size_t)n);
b.reserve((size_t)n);
for (int i = 0; i < n; i++) {
Vec3 da, db;
const size_t k = (size_t)i;
if (!to_direction(p.ax[k], p.ay[k], da)) continue;
if (!to_direction(p.bx[k], p.by[k], db)) continue;
a.push_back(da);
b.push_back(db);
}
const int m = (int)a.size();
if (m < 16) return -1.0f;
const float thr = 3.0f * rad_per_pixel();
uint32_t rng = 0x85ebca6bu;
std::vector<int> inliers, trial;
float best[9] = {1, 0, 0, 0, 1, 0, 0, 0, 1};
int best_in = 0;
for (int it = 0; it < kRansac; it++) {
const int i0 = (int)(next_random(rng) % (uint32_t)m);
const int i1 = (int)(next_random(rng) % (uint32_t)m);
float R[9];
if (i0 == i1 || !rotation_from_two(a[(size_t)i0], a[(size_t)i1],
b[(size_t)i0], b[(size_t)i1], R))
continue;
trial.clear();
for (int i = 0; i < m; i++)
if (angle_between(mul(R, a[(size_t)i]), b[(size_t)i]) < thr)
trial.push_back(i);
if ((int)trial.size() > best_in) {
best_in = (int)trial.size();
inliers = trial;
std::memcpy(best, R, sizeof best);
}
}
if (best_in < 8) return -1.0f;
rotation_from_many(a, b, inliers, best);
std::vector<float> res;
res.reserve((size_t)m);
const float clamp = 0.35f * std::max(o.out_fov, 0.25f);
for (int i = 0; i < m; i++)
res.push_back(std::min(
clamp, angle_between(mul(best, a[(size_t)i]), b[(size_t)i])));
return kParallaxWeight * percentile(res, 0.75f) /
std::max(o.out_fov, 0.25f);
}
float MotionTracker::Impl::planar_cost(const Pair& p) const {
const int n = (int)p.ax.size();
if (n < 16) return -1.0f;
const float thr = 0.004f * diag + 0.5f;
uint32_t rng = 0x9e3779b9u;
float best[6] = {1, 0, 0, 0, 1, 0};
int best_in = 0;
std::vector<int> inliers, trial;
for (int it = 0; it < kRansac; it++) {
int pick[3];
for (int k = 0; k < 3; k++)
pick[k] = (int)(next_random(rng) % (uint32_t)n);
float A[6];
if (!solve_affine(p.ax, p.ay, p.bx, p.by, pick, 3, A)) continue;
trial.clear();
for (int i = 0; i < n; i++) {
float u, v;
apply_affine(A, p.ax[(size_t)i], p.ay[(size_t)i], u, v);
const float ex = u - p.bx[(size_t)i], ey = v - p.by[(size_t)i];
if (ex * ex + ey * ey < thr * thr) trial.push_back(i);
}
if ((int)trial.size() > best_in) {
best_in = (int)trial.size();
inliers = trial;
std::memcpy(best, A, sizeof best);
}
}
if (best_in < 8) return -1.0f;
float A[6];
if (solve_affine(p.ax, p.ay, p.bx, p.by, inliers.data(),
(int)inliers.size(), A))
std::memcpy(best, A, sizeof best);
std::vector<float> res;
res.reserve((size_t)n);
const float clamp = 0.35f * diag;
for (int i = 0; i < n; i++) {
float u, v;
apply_affine(best, p.ax[(size_t)i], p.ay[(size_t)i], u, v);
const float ex = u - p.bx[(size_t)i], ey = v - p.by[(size_t)i];
res.push_back(std::min(clamp, std::sqrt(ex * ex + ey * ey)));
}
const float parallax = percentile(res, 0.75f) / diag;
// What the model says left the frame, and what came in: a camera backing
// away loses nothing and still sees a scene it has not seen.
const float w = (float)o.width, h = (float)o.height;
Poly moved;
for (int k = 0; k < 4; k++) {
const float x = (k == 1 || k == 2) ? w : 0.0f;
const float y = (k >= 2) ? h : 0.0f;
float u, v;
apply_affine(best, x, y, u, v);
moved.x[k] = u;
moved.y[k] = v;
}
moved.n = 4;
const float a_moved = poly_area(moved);
const float a_shared = poly_area(clip_to_frame(moved, w, h));
const float lost = a_moved > 0 ? 1.0f - a_shared / a_moved : 1.0f;
const float fresh = 1.0f - a_shared / (w * h);
const float coverage = std::min(1.0f, std::max(lost, fresh));
return coverage + kParallaxWeight * parallax;
}
float MotionTracker::Impl::step_cost(const Pair& p) const {
return sphere() ? sphere_cost(p) : planar_cost(p);
}
MotionTracker::MotionTracker(const MotionOptions& options) : impl_(new Impl{}) {
Impl& s = *impl_;
s.o = options;
if (s.o.circle_r <= 0) {
s.o.circle_x = 0.5f * (float)s.o.width;
s.o.circle_y = 0.5f * (float)s.o.height;
s.o.circle_r = 0.5f * (float)std::min(s.o.width, s.o.height);
}
s.diag = std::sqrt((float)(s.o.width * s.o.width + s.o.height * s.o.height));
// About 500 points: enough for a stable fit, few enough that a step costs
// a millisecond or two of one core.
const int cols = std::max(8, (int)std::lround(s.o.width / 18.0));
const int rows = std::max(8, (int)std::lround(s.o.height / 18.0));
for (int j = 0; j < rows; j++)
for (int i = 0; i < cols; i++) {
s.gx.push_back((i + 0.5f) / (float)cols * (float)s.o.width);
s.gy.push_back((j + 0.5f) / (float)rows * (float)s.o.height);
}
}
MotionTracker::~MotionTracker() = default;
void MotionTracker::track(const uint8_t* gray, int64_t index) {
Impl& s = *impl_;
if (s.o.width <= 0 || s.o.height <= 0 || s.done) return;
build_pyramid(gray, s.o.width, s.o.height, s.cur);
if (s.primed) {
// The points are independent and are most of the cost of a step, so
// they are split across the machine: the extraction this runs in front
// of has the GPU busy and the cores idle.
const size_t n = s.gx.size();
const int threads = std::max(
1, std::min((int)std::thread::hardware_concurrency(), (int)(n / 64)));
std::vector<Impl::Pair> part((size_t)threads);
const size_t span = (n + threads - 1) / (size_t)threads;
auto run = [&](int t) {
Impl::Pair& q = part[(size_t)t];
for (size_t k = t * span; k < std::min(n, (t + 1) * span); k++) {
float dx, dy;
if (!track_point(s.prev, s.cur, s.gx[k], s.gy[k], dx, dy)) continue;
q.ax.push_back(s.gx[k]);
q.ay.push_back(s.gy[k]);
q.bx.push_back(s.gx[k] + dx);
q.by.push_back(s.gy[k] + dy);
}
};
std::vector<std::thread> pool;
for (int t = 1; t < threads; t++) pool.emplace_back(run, t);
run(0);
for (std::thread& th : pool) th.join();
Impl::Pair p;
for (const Impl::Pair& q : part) {
p.ax.insert(p.ax.end(), q.ax.begin(), q.ax.end());
p.ay.insert(p.ay.end(), q.ay.begin(), q.ay.end());
p.bx.insert(p.bx.end(), q.bx.begin(), q.bx.end());
p.by.insert(p.by.end(), q.by.begin(), q.by.end());
}
const float c = s.step_cost(p);
if (c < 0) s.weak++;
s.cost.push_back(c);
s.end.push_back(index);
}
s.prev.level.swap(s.cur.level);
s.prev.w = s.cur.w;
s.prev.h = s.cur.h;
s.primed = true;
}
void MotionTracker::finish() {
Impl& s = *impl_;
if (s.done) return;
s.done = true;
// A step nothing could be tracked across is not a still one: the typical
// step is a better guess than zero, which would hoard frames there.
std::vector<float> good;
for (float c : s.cost)
if (c >= 0) good.push_back(c);
const float fill = good.empty() ? 0.0f : percentile(good, 0.5f);
for (float& c : s.cost)
if (c < 0) c = fill;
}
const std::vector<float>& MotionTracker::costs() const { return impl_->cost; }
const std::vector<int64_t>& MotionTracker::ends() const { return impl_->end; }
int MotionTracker::weak_steps() const { return impl_->weak; }
float motion_out_fov(const std::vector<Pano360View>& views) {
if (views.empty()) return 1.5708f;
const Pano360View& v = views[0];
if (v.fov <= 0.0f) return 3.14159f; // equirectangular
const double focal = v.width * 0.5 / std::tan(v.fov * kPi / 360.0);
const double diag =
std::sqrt((double)v.width * v.width + (double)v.height * v.height);
return (float)(2.0 * std::atan(0.5 * diag / focal));
}
void motion_frame_size(MotionView view, int src_w, int src_h, int& w, int& h) {
w = h = 0;
if (src_w <= 0 || src_h <= 0) return;
const double want = view == MotionView::Planar ? 320.0 * 240.0 : 640.0 * 480.0;
const double s = std::min(1.0, std::sqrt(want / ((double)src_w * src_h)));
w = std::max(64, 2 * (int)std::lround(src_w * s * 0.5));
h = std::max(64, 2 * (int)std::lround(src_h * s * 0.5));
}
// ---------------------------------------------------------------------------
// Planning
// ---------------------------------------------------------------------------
std::vector<int64_t> plan_by_motion(const std::vector<float>& cost,
const std::vector<int64_t>& ends,
int64_t frames, int skip, int window,
float range, int max_frames) {
std::vector<int64_t> out;
if (cost.empty() || cost.size() != ends.size() || skip < 1) return out;
if (range < 1.0f) range = 1.0f;
std::vector<double> sum(cost.size());
double total = 0;
for (size_t i = 0; i < cost.size(); i++) {
total += std::max(0.0f, cost[i]);
sum[i] = total;
}
int64_t want = std::max<int64_t>(1, frames / skip);
if (max_frames > 0) want = std::min<int64_t>(want, max_frames);
// Never closer than one sharpness window: two windows that overlap can
// choose the same frame, and one of the two kept frames then vanishes.
const int64_t min_gap = std::max<int64_t>(
std::max(1, window), (int64_t)((double)skip / range));
const int64_t max_gap =
std::max<int64_t>(min_gap, (int64_t)((double)skip * range));
auto walk = [&](double step) {
std::vector<int64_t> got;
int64_t last = -1;
double last_sum = 0;
for (size_t i = 0; i < cost.size(); i++) {
const int64_t at = ends[i];
if (at < window - 1) continue;
size_t pick = i;
if (last >= 0) {
if (at - last < min_gap) continue;
if (sum[i] - last_sum < step && at - last < max_gap) continue;
// The step falls BETWEEN two samples, and always taking the one
// past it lands the whole plan late.
if (i > 0 && sum[i] - last_sum > step && ends[i - 1] > last &&
ends[i - 1] - last >= min_gap && ends[i - 1] >= window - 1 &&
(sum[i] - last_sum) - step > step - (sum[i - 1] - last_sum))
pick = i - 1;
}
got.push_back(ends[pick]);
last = ends[pick];
last_sum = sum[pick];
i = pick; // the sample stepped over is still the next candidate
if (max_frames > 0 && (int)got.size() >= max_frames) break;
}
return got;
};
// Bisected rather than total/want, because a burst of motion swallows
// several steps' budget in one sample and spends one frame of it. The floor
// is what stops a still capture being answered with every frame of itself.
const double floor_step = 0.01;
double lo = floor_step, hi = std::max(floor_step * 2.0, total);
out = walk(lo);
if ((int64_t)out.size() > want) {
for (int it = 0; it < 32; it++) {
const double mid = 0.5 * (lo + hi);
std::vector<int64_t> got = walk(mid);
if ((int64_t)got.size() > want) {
lo = mid;
} else {
hi = mid;
out = std::move(got);
}
}
}
return out;
}
} // namespace app
+89
View File
@@ -0,0 +1,89 @@
#pragma once
// FrameMotion -- how fast a video's view is changing, and the frame spacing
// that follows from it.
//
// A grid of points tracked between small grey frames, one global model fitted
// to them -- a 2D affine, or a rotation of the sphere where the capture sees
// all of it -- and two numbers out: how much of the view left it, and how much
// of the flow that model could not explain. A capture that only turns on the
// spot scores nothing on the second, which is what tells a pan apart from a
// walk past something close. docs/datasets.md weighs them.
#include "app/Pano360.h"
#include <cstdint>
#include <memory>
#include <vector>
namespace app {
// What the grey frames are pictures of, which is what decides whether turning
// the camera costs anything: a full-sphere capture keeps every direction it
// had, a rectilinear one does not.
enum class MotionView { Planar, Eac360, Fisheye };
struct MotionOptions {
MotionView view = MotionView::Planar;
int width = 0, height = 0; // of the grey frames handed to track()
// Eac360: the packing, in SOURCE pixels, of the ONE track the grey frames
// come from (its three faces are half the sphere, enough to fit a rotation).
Eac360Layout eac;
// Fisheye: the image circle in grey-frame pixels (0 = the inscribed one).
// Every 360 lens is a little over 180 and none of them agree; a few degrees
// of error leaves less residual than the tracking floor already does.
float circle_x = 0, circle_y = 0, circle_r = 0;
float circle_fov = 3.4034f; // full angle across the circle, radians
// Radians across the diagonal of what the dataset will hold, which is what
// an angular residual is a fraction of.
float out_fov = 1.5708f;
};
// Tracks one stream of grey frames. Costs are readable only after the last
// track(): a fisheye's field of view is fitted from the first few steps and
// their costs are revised once it is.
class MotionTracker {
public:
explicit MotionTracker(const MotionOptions& options);
~MotionTracker();
MotionTracker(const MotionTracker&) = delete;
MotionTracker& operator=(const MotionTracker&) = delete;
// `gray` is width*height bytes; `index` is the source frame it came from.
// The first call only primes the reference.
void track(const uint8_t* gray, int64_t index);
// Fills in the steps nothing could be tracked across. Call it once, after
// the last track().
void finish();
// One entry per step, in order: the view change across it, and the source
// frame it ends at. Both are final only after finish().
const std::vector<float>& costs() const;
const std::vector<int64_t>& ends() const;
// Steps whose flow was too weak to fit a model to, for the log.
int weak_steps() const;
private:
struct Impl;
std::unique_ptr<Impl> impl_;
};
// The angle across the diagonal of the views a 360 plan writes, which is what
// an angular residual is a fraction of. 90 degrees without a plan.
float motion_out_fov(const std::vector<Pano360View>& views);
// The grey frames to hand the tracker, for a source of this size. A full
// sphere spans three times the angle a rectilinear frame does, so it is given
// the pixels to match or its own tracking noise drowns the parallax.
void motion_frame_size(MotionView view, int src_w, int src_h, int& w, int& h);
// Which source frames a run should end its sharpness windows at, so that kept
// frames differ by view rather than by time. `skip` is the fixed schedule's
// spacing, and the rate stays within `range` of it.
std::vector<int64_t> plan_by_motion(const std::vector<float>& cost,
const std::vector<int64_t>& ends,
int64_t frames, int skip, int window,
float range, int max_frames);
} // namespace app
+40 -3
View File
@@ -86,15 +86,33 @@ void dir_to_cell(float x, float y, float z, int& cell, float& u, float& v) {
}
}
// dir_to_cell run backwards, cell by cell: the face axis is +-1 and the two
// in-cell coordinates are what the table above put where.
void cell_to_dir(int cell, float u, float v, float d[3]) {
switch (cell) {
case 0: d[0] = -1; d[1] = v; d[2] = u; break;
case 1: d[0] = u; d[1] = v; d[2] = 1; break;
case 2: d[0] = 1; d[1] = v; d[2] = -u; break;
case 3: d[0] = -v; d[1] = 1; d[2] = -u; break;
case 4: d[0] = -v; d[1] = -u; d[2] = -1; break;
default: d[0] = -v; d[1] = -1; d[2] = u; break;
}
const float n = std::sqrt(d[0] * d[0] + d[1] * d[1] + d[2] * d[2]);
d[0] /= n;
d[1] /= n;
d[2] /= n;
}
} // namespace
bool eac360_detect(int tracks, int width, int height, Eac360Layout& out) {
out = Eac360Layout{};
if (tracks != 2 || width <= 0 || height <= 0) return false;
const int strips = width - 3 * height;
// 64 px in both recording modes. A generous ceiling still rejects every
// other two-track file: an Insta360 .insv is two square-ish fisheyes.
if (strips < 0 || strips % 2 != 0 || strips > 128) return false;
// 2x32 px at 5.6K and 3K, 2x96 at 8K. A ceiling that scales with the face
// still rejects every other two-track file: an Insta360 .insv is two
// square-ish fisheyes.
if (strips < 0 || strips % 2 != 0 || strips > height / 8) return false;
out.track_w = width;
out.track_h = height;
out.face = height;
@@ -109,6 +127,25 @@ std::vector<Eac360Slice> eac360_slices(const Eac360Layout& l) {
{2 * l.face + half + 2 * l.strip, 2 * l.face + half, half}};
}
bool eac360_direction(const Eac360Layout& l, int row, float x, float y,
float dir[3]) {
if (!l.valid() || row < 0 || row > 1 || y < 0 || y >= (float)l.track_h)
return false;
int dst_x = -1;
for (const Eac360Slice& s : eac360_slices(l))
if (x >= (float)s.src_x && x < (float)(s.src_x + s.width))
dst_x = (int)(s.dst_x + (x - (float)s.src_x));
if (dst_x < 0) return false;
const float face = (float)l.face;
const int col = std::min(2, (int)(dst_x / l.face));
const float uf = 2.0f * ((float)dst_x - col * face) / face - 1.0f;
const float vf = 2.0f * y / face - 1.0f;
cell_to_dir(row * 3 + col, (float)std::tan(kPi / 4 * uf),
(float)std::tan(kPi / 4 * vf), dir);
return true;
}
int pano360_default_size(const Eac360Layout& l, const Pano360Options& o) {
if (!l.valid()) return 0;
// Four faces around a panorama, so a pixel spans an EAC pixel's angle. A
+7 -1
View File
@@ -29,7 +29,7 @@ struct Eac360Layout {
};
// Two equal video tracks whose width is three square faces plus two small
// strips. Both recording modes answer this: 4096x1344 (5.6K) and 2272x736.
// strips: 4096x1344 (5.6K), 2272x736 (3K), 5952x1920 (8K).
bool eac360_detect(int tracks, int width, int height, Eac360Layout& out);
// One column range of a track, and where it lands in the canvas row. Each
@@ -40,6 +40,12 @@ struct Eac360Slice {
};
std::vector<Eac360Slice> eac360_slices(const Eac360Layout& l);
// A track pixel's viewing direction, unit, in the frame Pano360View::rot maps
// into. False inside an overlap strip -- the one place the packing stores two
// answers -- and outside the track.
bool eac360_direction(const Eac360Layout& l, int row, float x, float y,
float dir[3]);
enum class Pano360Mode { Off, Faces, Equirect };
struct Pano360Options {
+3 -2
View File
@@ -99,7 +99,7 @@ GUI (Basic Options) for remote monitoring.
| `gui/MaskSettings.h` | The mask prompt and its thresholds, split out of `SegmentPanel.h` so a preset module does not depend on a UI panel. |
| `gui/BatchProcess.h/.cpp` | **Batch processing**: a list of rows, each of which can build a dataset, train it any number of times, and mesh what came out -- and none of those three has to exist when the queue is started. A row carries the videos or photo folders to build from, the dataset folder (which the Dataset stage WRITES and the Train stage READS, so the two cannot disagree), a dataset preset, a list of `BatchRun`s and a `BatchMeshOptions`. A **run** is one training pass: its preset, and text overrides of that preset's **max splats / SH degree / steps** -- text because "unset" and "0" are different answers and 0 is a legal `--sh-degree` -- plus whether the row's Mesh stage covers it, which is what lets one row train a large model for its appearance and a cheap one beside it for the mesh. The **mesh options** are a preset plus the colors and formats to write over what it says; empty means "leave the preset alone" rather than "write nothing". `batch_plan()` expands the rows into the tasks they will run as -- one per stage per run, and one mesh per run that asked for it -- which is both what the screen previews before the start and what the driver walks. `batch_progress()` turns the same list into the two bars and the estimate the screens draw, from what the finished tasks of each stage actually took. The list is data, not a runner: `GuiApp::advance_batch()` launches a task, waits for the runner that owns that stage to go idle, records what happened and launches the next -- driving the live screens rather than being a fourth implementation of what `SfmRunner`, `TrainRunner` and `MeshRunner` already are. A task that fails is an ordinary transition: it is recorded, the tasks downstream of it in the same row are passed over (what they needed was never produced), and the next row still gets its turn. `batch_check_row()` is the "warn before anything starts" half, and it is what makes a weekend-long queue worth setting up: an input that is not there, a preset file that has gone missing, a segmentation or geometry checkpoint not downloaded (a batch cannot stop to accept a licence), masking with no prompt (a batch cannot be prompted with clicks), a flag `train_config_unsupported()` refuses, a folder that already holds a reconstruction the run would ADD to rather than rebuild, two rows building into one folder, a row training a dataset a LATER row builds, a preset made for video pointed at photographs, a meshing stage with no run ticked for it, and colors and formats with no pair between them. Nothing waits for a button: the check runs on every edit and on a timer, so a preset file deleted while the screen is open is reported where it happened. Fatal issues block the start (with a "skip the bad rows" way past); the rest are said out loud. The list is kept in `<config_dir>/batch.json` -- losing a queue to a crash three hours in is the one failure a user cannot recover from by trying again -- and a file from the train-only era still loads. A finished queue can **run a command** (`batch_command` in `gui.conf`, so it belongs to the machine rather than to the rows): `kBatchMessageToken` -- `{message}` -- is replaced by a one-line summary of how it went and always arrives as a single argument, braces because `<>`, `%..%` and `$..` are all punctuation some shell would eat. It runs however the queue ended, and the Test button beside the field sends a test message instead, which is the only honest way to find out whether a notifier works before leaving a queue overnight. The field is a box that grows with what is in it (one line while empty, up to twelve), because the real command is a `curl` with a JSON body that nobody writes on one line -- `gui.conf` is one setting per line, so the breaks are escaped on the way in and out. |
| `gui/ColmapRunner.h/.cpp` | images/video → COLMAP dataset. Requires **COLMAP >= 4.x** (version probed from `colmap help`; 3.x-era `use_gpu` flags are gone — flags follow `reference/scripts/run_colmap.bash`): feature_extractor (**SIFT or ALIKED** via `FeatureExtraction.type`; camera model = any the parser supports; single camera / per-folder / per-image; optional **initial focal length** — `init_focal_factor` × width probed via `stbi_info`, composed into `ImageReader.camera_params` with centered principal point + zero distortion, or a raw `camera_params` string) → **explicitly chosen** exhaustive / sequential / vocab-tree matcher (no "auto"; the GUI presets sequential for video, exhaustive for photos). Sequential supports overlap + **quadratic overlap** (on by default) + **loop closure** (`SequentialMatching.loop_detection` via the vocab tree — SIFT only; the tree is auto-found near the workspace/cache or downloaded via curl into `~/.cache/spirula-studio`). **LightGlue** matching (`FeatureMatching.type` `SIFT_LIGHTGLUE`/`ALIKED_LIGHTGLUE`, default for ALIKED) → mapper (`ba_use_gpu` — **forced OFF for fisheye models**, COLMAP's GPU BA doesn't support them; `Mapper.ba_refine_extra_params 0` for perspective models by default — distortion held fixed for stability and recovered in the final BA, per run_colmap.bash's advice; `min_num_matches`) → best-effort **model_merger** when the mapper splits (kept only if the merged model registers more images; written as the next `sparse/<N>` — on a real X5 capture this fused 86+39 partials into 116/118 frames) → optional bundle_adjuster refinement **on the largest/merged model** (was hardcoded `sparse/0`), with a **verify-and-revert guard**: mean reprojection error before/after via `model_analyzer`, refinement discarded when worse/non-finite (releasing pp + 8 thin-prism coefficients on a ~200° fisheye reliably diverges to ~1e150 px — pp additionally stays fixed for fisheye). `.insv` preset: THIN_PRISM_FISHEYE, per-folder cameras, focal factor 0.269 (Insta360 X5: fx=fy≈0.269·width), **exhaustive** matcher (the two lens tracks are concatenated, so sequential misses cross-lens pairs: 116/118 exhaustive vs 68/118 sequential+loop on the same capture). **Resume** (`ColmapJob::resume`, GUI checkbox appears when the workspace holds a previous run): extracted frames are kept per track, mask.py resumes on its own, COLMAP skips features/matches already in database.db, and existing `sparse/<N>` models skip the mapper (the mapper only writes models on completion, so partials are never trusted); with resume off a non-empty workspace is refused rather than mixed into. A measured resume replay of a full .insv run took 39 s vs 246 s. **Repetitive-scene knobs** (Advanced → "Repetitive scenes", with an Off/Low/Medium/High preset combo; all 0 = default): `SiftMatching.max_ratio`, `TwoViewGeometry.min_num_inliers`, `Mapper.abs_pose_min_num_inliers` / `abs_pose_min_inlier_ratio` / `abs_pose_max_error` — stricter matching + registration so similar-looking rooms don't weld together. The GUI's default workspace auto-suffixes `_2`, `_3`, ... instead of pointing at an existing non-empty folder. Optional **AI masking**: the CMake-embedded `reference/scripts/mask.py` runs via external Python (lang-segment-anything / SAM-3), masks feed `ImageReader.mask_path` + the trainer's masks dir; missing-package output is detected and surfaced with install instructions. Photo-folder inputs are indexed **in place**, recursively (no copy) — the absolute image dir is handed to the GUI in-memory for the immediate open; no marker file is written, so on later re-opens set `data.image_dir` in the dataparser options (video datasets use the default `images/`). Multi-track `.insv` videos split into `images/cam<N>/` + per-folder cameras. The prep stage is shared, so a COLMAP run can take several inputs too — but `feature_extractor` fits ONE camera model to the whole run, so the per-input lens models only mean anything on the built-in path (the panel says so). The mapper may emit several partial models — the trainers auto-pick the largest (below). |
| `gui/FrameSelect.h/.cpp` | Blur-aware video frame selection (multithreaded C++ port of `reference/scripts/extract_frames.py`): ffmpeg extracts candidates at (target fps x sharpness window) with `-nostdin` (**without it ffmpeg can hang reading stdin** — the "stuck at ffmpeg" bug), then the sharpest per window is kept (variance of 3x3 Laplacian on mean-subtracted 512^2 gray, scored across all cores). |
| `gui/FrameSelect.h/.cpp` | Blur-aware video frame selection (multithreaded C++ port of `reference/scripts/extract_frames.py`): ffmpeg extracts candidates at (target fps x sharpness window) with `-nostdin` (**without it ffmpeg can hang reading stdin** — the "stuck at ffmpeg" bug), then the sharpest per window is kept (variance of 3x3 Laplacian on mean-subtracted 512^2 gray, scored across all cores). With the adaptive rate on it extracts at (target fps x max(window, spread)) instead and hands the candidates to `FrameMotion`, which decides which of them to keep -- so the fallback adapts without a second decode. |
| `gui/FileDialog.h/.cpp` | Built-in ImGui file/folder browser (no native-dialog dep; works over WSLg/remote X). File mode takes an optional **multi-select** (`open(..., multi_select=true)`, `results()`): clicking a file toggles it, so a dataset can be built from several clips in one pick. |
| `gui/DatasetPrep.h/.cpp` | Everything between "the user picked an input" and "there is an image directory ready for SfM", shared by both reconstruction paths: frame extraction (in-process VK video decode → ffmpeg fallback), sharpest-frame selection, `.insv` track split, AI masking, and the resume rules. A job takes a **list of inputs** (`PrepInput`: path, video-or-folder, the `images/<subdir>` its frames go to, the masks that came with it, and the lens + focal factor that belong to it). **A picked photo folder is resolved by the project's own layout conventions** (`resolve_photo_folder`, matching `spirula sfm auto`'s probing and the dataparsers' `mask_dir = "masks"`): `<picked>/images` is the folder to index when it exists, and the `masks/` beside it holds masks that are **already made** -- those are used as they are and that input is never AI-masked (generating over them would write into a folder we were only asked to read). A folder named `masks` is never an input of its own: picked or dropped, it attaches to the input whose images it sits beside. **Every image walk follows directory symlinks** -- a prepared capture whose `images/`+`masks/` are links into the raw one is an ordinary layout, and the default iterator returns nothing for it -- and never descends into a `masks/` nested under the images, which would otherwise double the dataset with PNGs that are not views. One folder of photos is still read where it is; anything else is gathered into the dataset's own `images/`, one sub-folder per input (a multi-track video adds `cam0/`, `cam1/` under that), which is what makes the inputs separate cameras. `camera_subfolders` reports the folders that hold images **directly**, at any depth and including the root, so a capture handed over as `1/cam0`, `1/cam1`, `2/cam0` ... is as many camera groups as `--camera-mode folder` will make of it; it is depth- and count-bounded because each one becomes a row in the dataset panel. Photo folders in a multi-input job are **hard-linked** into place (masks too, so the two trees keep mirroring), copied when the filesystem refuses. `probe_workspace` reports what the output folder already holds, split into what a run can **reuse** (extracted frames / features / matches / masks -> the resume checkbox) and what it would **replace** (`sparse/` -> a warning): an existing reconstruction is never treated as resumable, since a folder holding one opens in the trainer when dropped, so anything that gets to this screen with a `sparse/` in it is someone else's dataset or a finished one. The input's own images and masks never count as leftovers -- which matters because the output folder for a capture that already has `images/` **is that folder**: `sparse/` belongs beside `images/` and `masks/`, which is exactly where every parser looks. Masks mirror the image tree and are generated **per input**, so the tracker's memory bank never crosses from one capture into the next; a clicked object prompts only the input it was drawn on (`MaskClick::source`), and a clicks-only job whose inputs are not all prompted is **refused** rather than run -- a dataset where three of four camera folders kept the subject reads as a masking run that worked. |
| `gui/SfmRunner.h/.cpp` | images/video → dataset with the built-in SfM, by re-running this executable as `spirula sfm auto` (see SfmRunner.h for why it is a child process). Runs `DatasetPrep` first, then maps the panel onto flags. Per-input lenses become **`--camera-model DIR=MODEL` / `--focal DIR=PX`** overrides, `DIR` being the group's path under `images/` — the input's sub-folder, or a camera folder inside it at any depth (`1/cam0`), which is exactly how `--camera-mode folder` groups. A video input keeps one row: a dual-lens file's `cam0/`+`cam1/` sit under its prefix and are still two camera groups. The rows are built once (`camera_groups` in DatasetPrep), so what the panel draws and what the child is handed cannot drift, and an **empty model on a row means "same as above"** — a dozen clips off one camera are one choice, and the overrides come out fully resolved. The focal is carried as a fraction of the image width and resolved to pixels once the frames exist (`first_image_dims`); an explicit focal typed into Advanced wins for a lone input. Masks reach the run as `--masks <dir>` (or `--no-masks`, so a stale `masks/` beside the images is never picked up silently); `PrepResult::mask_dir_cfg` and `image_dir_cfg` are handed to `GuiApp::open_dataset` in memory, so photos read where they are keep both their image and mask folders when the dataset opens in the trainer. Camera sharing switches to per-folder on its own when `images/` came out with sub-folders. Exit code 3 = reconstructed but partial (reported, not failed). |
@@ -138,7 +138,8 @@ extraction (see below).
| `Json.h` | Minimal dependency-free JSON parser (handles Python `Infinity`/`NaN`). |
| `Xml.h` | Minimal dependency-free XML parser (ElementTree subset: `attr`/`find`/`findall`/recursive `iter`; comments/PI/CDATA/DOCTYPE skipped, standard entities). Metashape nests `<camera>` in `<group>` and `<camera_ids>` in `<partition>` — the recursive `iter` is load-bearing. |
| `Knn.h` | Exact kd-tree kNN (multi-threaded queries) for `seed_splats` scale init; replaced the hash-grid approx that degenerated to O(N²) on SfM outliers. |
| `Pano360.h/.cpp` | A 360 camera's own frame layout and the views a dataset wants out of it: the GoPro MAX `.360` EAC packing (two tracks, one 3x2 cubemap, a 32 px lens-seam overlap at the centre of each track's side faces), the plan that turns it into ten seam-free perspective views (five per lens, so that no view spans both) or one 2:1 panorama, and the resampler that applies it. **Both decode paths go through this one implementation** -- the in-process decoder hands it two decoded tracks, and the ffmpeg fallback runs a filter graph generated from the same layout that only crops and stacks. ffmpeg is never asked to warp: its own `v360=eac` insets every face by 2 px, which puts a 4 px step across the seam between the two tracks. The canvases the two paths build are bit-identical. `docs/datasets.md` records where the layout numbers were measured. |
| `FrameMotion.h/.cpp` | How fast a video's view is changing, and the frame spacing that follows. A grid of ~500 points tracked between small grey frames by pyramidal Lucas-Kanade, one global model fitted by RANSAC -- a 2D affine for an ordinary video, a rotation of the SPHERE for a `.360` or a dual fisheye -- and two numbers out: how much of the view left it (0 on a sphere, which is what makes turning a 360 camera free) and how much of the flow the model could not explain (parallax, which is what a walk past something close produces). `plan_by_motion` then spaces the kept frames at equal cumulative cost, bisecting for the step that lands on the wanted count. Portable C++, no device: it runs in front of the built-in decoder (`FrameExtract`, from a GPU-reduced grey frame) and behind the ffmpeg one (`gui/FrameSelect`, from the candidate JPEGs). `docs/datasets.md` "Frames out of a video" has the weights and why. |
| `Pano360.h/.cpp` | A 360 camera's own frame layout and the views a dataset wants out of it: the GoPro MAX `.360` EAC packing (two tracks, one 3x2 cubemap, a lens-seam overlap -- 32 px at 5.6K and 3K, 96 px at 8K -- at the centre of each track's side faces), the plan that turns it into ten seam-free perspective views (five per lens, so that no view spans both) or one 2:1 panorama, and the resampler that applies it. **Both decode paths go through this one implementation** -- the in-process decoder hands it two decoded tracks, and the ffmpeg fallback runs a filter graph generated from the same layout that only crops and stacks. ffmpeg is never asked to warp: its own `v360=eac` insets every face by 2 px, which puts a 4 px step across the seam between the two tracks. The canvases the two paths build are bit-identical. `docs/datasets.md` records where the layout numbers were measured. |
| `WriterPool.h` | Bounded-queue worker threads that JPEG/PNG-encode and write frames and masks off the calling thread. Used by `FrameExtract`, `spirula sam track` and the GUI's folder-masking loop. Encoding a 1080p mask through stb's deflate is ~75 ms — a third of a SAM 2.1 Tiny frame — and none of it needs the GPU, so a caller that writes inline sets the frame rate with zlib. The queue bound is what keeps a slow disk applying back-pressure instead of growing until memory runs out. |
| `HttpServer.h/.cpp` | Minimal HTTP/1.0 GET server (POSIX sockets; winsock shim compiles but untested). Serial request handling — parity with Python's non-threading `HTTPServer`. |
| `Viewer.h/.cpp` | Web-viewer server: latest-wins render worker, the viewer buffer subset, `engine_blit_view` GPU annotation/colormap, stb JPEG encode. Serves the **unchanged** `viewer.html` (embedded at configure time via CMake hex; `SS_VIEWER_HTML=<path>` env overrides for dev). `/pick?px=&py=&<camera params>` returns the 3D point under a pixel as JSON for viewer.html's double-click centering (the client treats non-OK responses as a no-op). |
+9
View File
@@ -83,6 +83,8 @@ void usage() {
help_row(" --scale <f>", H::xh_scale);
help_row(" --track <i>", H::xh_track);
help_row(" --sync", H::xh_sync);
help_row(" --adaptive", H::xh_adaptive);
help_row(" --adaptive-range <f>", H::xh_adaptive_range);
help_row(" --threads <n>", H::xh_threads);
std::fprintf(stderr, "\n%s\n", H::xh_360_section.get());
@@ -118,6 +120,8 @@ struct Options {
float scale = 1.0f;
int track = -1;
bool sync = false;
bool adaptive = false;
float adaptive_range = 4.0f;
int threads = 0;
std::string pano_mode = "faces";
@@ -157,6 +161,9 @@ bool parse_args(int argc, char** argv, Options& o) {
else if (a == "--scale") o.scale = std::strtof(next("--scale"), nullptr);
else if (a == "--track") o.track = std::atoi(next("--track"));
else if (a == "--sync") o.sync = true;
else if (a == "--adaptive") o.adaptive = true;
else if (a == "--adaptive-range")
o.adaptive_range = std::strtof(next("--adaptive-range"), nullptr);
else if (a == "--threads") o.threads = std::atoi(next("--threads"));
else if (a == "--360") o.pano_mode = next("--360");
else if (a == "--360-size") o.pano.size = std::atoi(next("--360-size"));
@@ -252,6 +259,8 @@ int sam_cli_extract(int argc, char** argv) {
job.scale = o.scale;
job.track = o.track;
job.sync_tracks = o.sync;
job.adaptive = o.adaptive;
job.adaptive_range = o.adaptive_range;
job.threads = o.threads;
job.write_overlay = o.overlay;
// A 360 file is recognised by its packing, not by its name, and only then
+6
View File
@@ -246,12 +246,16 @@ void ColmapRunner::take_reconstruction(ColmapJob& job) {
const std::string workspace = job.workspace;
const bool resume = job.resume;
const float fps = job.video_fps;
const bool adaptive = job.adaptive_fps;
const float range = job.adaptive_range;
const int sharp = job.sharp_window, maxf = job.max_frames;
job = _live;
job.inputs = inputs;
job.workspace = workspace;
job.resume = resume;
job.video_fps = fps;
job.adaptive_fps = adaptive;
job.adaptive_range = range;
job.sharp_window = sharp;
job.max_frames = maxf;
}
@@ -504,6 +508,8 @@ void ColmapRunner::run(ColmapJob job) {
pj.redo_masks = job.redo_masks;
pj.photo_import = job.photo_import;
pj.video_fps = job.video_fps;
pj.adaptive_fps = job.adaptive_fps;
pj.adaptive_range = job.adaptive_range;
pj.sharp_window = job.sharp_window;
pj.pano = job.pano;
pj.max_frames = job.max_frames;
+2
View File
@@ -106,6 +106,8 @@ struct ColmapJob {
// Video extraction
float video_fps = 2.0f; // kept frames per second
bool adaptive_fps = false; // see PrepJob
float adaptive_range = 4.0f;
int sharp_window = 3; // pick sharpest of N candidates (1 = off)
app::Pano360Options pano; // see PrepJob
int max_frames = 100000;
+62 -34
View File
@@ -479,18 +479,18 @@ private:
Clock::time_point _start, _last;
};
// One frame is written every this many source frames, from the run's kept
// frame rate and what the container says it holds.
int frame_skip(const PrepJob& job, double src_fps) {
return std::max(1, (int)std::lround(src_fps / std::max(job.video_fps, 0.01f)));
// One frame is written every this many source frames, from the kept frame rate
// this input asked for and what the container says it holds.
int frame_skip(float fps, double src_fps) {
return std::max(1, (int)std::lround(src_fps / std::max(fps, 0.01f)));
}
// What a video is expected to yield, for the step's bar. The extraction loop
// stops on the real end of stream either way.
int64_t expected_frames(const PrepJob& job, double src_fps, int64_t src_frames,
int tracks) {
int64_t expected_frames(const PrepJob& job, float fps, double src_fps,
int64_t src_frames, int tracks) {
int64_t expect = src_frames > 0
? (src_frames / frame_skip(job, src_fps)) * (int64_t)tracks
? (src_frames / frame_skip(fps, src_fps)) * (int64_t)tracks
: 0;
if (job.max_frames > 0 &&
(expect == 0 || expect > (int64_t)job.max_frames * tracks))
@@ -540,9 +540,12 @@ bool ffmpeg_probe_video(const std::string& ffmpeg_exe, const std::string& path,
out.duration = hh * 3600.0 + mm * 60.0 + ss;
}
// "... 1920x1080, 19938 kb/s, 30.01 fps, 30 tbr, ..."
// A cover picture is a video stream to ffmpeg and is not a
// lens; counting it gives a DJI .osv three of them.
const size_t f = line.find(" fps");
if (f == std::string::npos ||
line.find("Video:") == std::string::npos)
line.find("Video:") == std::string::npos ||
line.find("(attached pic)") != std::string::npos)
return;
// The frame size off the same line. The 16-pixel floor is
// what rejects the fourcc ("0x31637661"), which is also
@@ -989,7 +992,8 @@ int64_t DatasetPrep::estimate_frames(const PrepJob& job, const PrepInput& in,
std::string err;
video::VideoProbe probe;
if (video::probe_video(in.path, probe, err) && probe.tracks > 0)
return expected_frames(job, probe.fps > 1.0 ? probe.fps : 30.0,
return expected_frames(job, input_fps(job, in),
probe.fps > 1.0 ? probe.fps : 30.0,
probe.frame_count,
per_frame > 0 ? per_frame : probe.tracks);
}
@@ -997,7 +1001,7 @@ int64_t DatasetPrep::estimate_frames(const PrepJob& job, const PrepInput& in,
VideoFacts facts;
if (ffmpeg_probe_video(job.ffmpeg_exe, in.path, facts, _cancel) &&
facts.fps > 1.0)
return expected_frames(job, facts.fps, facts.frames,
return expected_frames(job, input_fps(job, in), facts.fps, facts.frames,
per_frame > 0 ? per_frame
: (is_dual_fisheye_path(in.path) ? 2 : 1));
return 0;
@@ -1294,7 +1298,7 @@ bool DatasetPrep::extract_video(const PrepJob& job, const PrepInput& in,
const bool ok = in.eac360.valid() && job.pano.mode != app::Pano360Mode::Off
? extract_360_ffmpeg(job, in, images, out, error)
: extract_video_ffmpeg(job, in, images, out, error);
if (ok) out.captures.push_back({in.subdir, in.path, (double)job.video_fps});
if (ok) out.captures.push_back({in.subdir, in.path, (double)input_fps(job, in)});
return ok;
}
@@ -1328,7 +1332,7 @@ bool DatasetPrep::extract_video_builtin(const PrepJob& job, const PrepInput& in,
const double src_fps = probe.fps > 1.0 ? probe.fps : 30.0;
const int window = std::max(job.sharp_window, 1);
const int skip = frame_skip(job, src_fps);
const int skip = frame_skip(input_fps(job, in), src_fps);
app::FrameExtractJob fx;
fx.input = in.path;
@@ -1337,6 +1341,8 @@ bool DatasetPrep::extract_video_builtin(const PrepJob& job, const PrepInput& in,
fx.keep = window > 1 ? window : 0;
fx.max_frames = job.max_frames;
fx.sync_tracks = job.sync_tracks;
fx.adaptive = job.adaptive_fps;
fx.adaptive_range = job.adaptive_range;
fx.auto_rotate = job.auto_rotate;
fx.quality = 95;
if (!views.empty()) {
@@ -1391,25 +1397,27 @@ bool DatasetPrep::extract_video_ffmpeg(const PrepJob& job, const PrepInput& in,
}
log(fmt(lmsg::video_input, {in.path}), /*detail=*/false);
// Multi-track videos (Insta360 .insv): one folder per track, one camera
// per folder, as in reference/scripts/extract_frames.py.
std::vector<int> streams = {0};
if (is_dual_fisheye_path(in.path)) {
std::vector<int> found;
run_process({"ffprobe", "-v", "error", "-select_streams", "v",
"-show_entries", "stream=index", "-of", "csv=p=0",
in.path}, "",
[&](const std::string& l) {
try { found.push_back(std::stoi(l)); } catch (...) {}
}, _cancel);
if (found.size() > 1) streams = found;
}
if (streams.size() > 1) out.per_folder_cameras = true;
// Multi-track videos (an Insta360 .insv, a DJI .osv): one folder per track,
// one camera per folder. Counted off the stream table rather than ffprobe's
// stream list, which calls an attached cover picture a video track.
VideoFacts facts;
ffmpeg_probe_video(job.ffmpeg_exe, in.path, facts, _cancel);
size_t streams = 1;
if (is_dual_fisheye_path(in.path) && facts.tracks.size() > 1)
streams = facts.tracks.size();
if (streams > 1) out.per_folder_cameras = true;
const int window = std::max(job.sharp_window, 1);
for (size_t tr = 0; tr < streams.size(); tr++) {
// Adaptive selection picks from the candidates, so there have to be enough
// of them for the fastest rate it may ask for.
const int group = job.adaptive_fps
? std::max(window, (int)std::ceil(job.adaptive_range))
: window;
const bool fisheye = streams > 1 && !facts.tracks.empty() &&
facts.tracks[0].first == facts.tracks[0].second;
for (size_t tr = 0; tr < streams; tr++) {
std::string track_path = in.path;
const fs::path out_dir = streams.size() > 1
const fs::path out_dir = streams > 1
? fs::path(images) / ("cam" + std::to_string(tr))
: fs::path(images);
if (job.resume && !job.redo_frames &&
@@ -1418,7 +1426,7 @@ bool DatasetPrep::extract_video_ffmpeg(const PrepJob& job, const PrepInput& in,
/*detail=*/false);
continue;
}
if (streams.size() > 1) {
if (streams > 1) {
enter(Stage::Frames, fmt(lmsg::stage_split_track, {(long long)tr}));
const fs::path tmp_track =
ws / ("track_cam" + std::to_string(tr) + ".mp4");
@@ -1440,7 +1448,7 @@ bool DatasetPrep::extract_video_ffmpeg(const PrepJob& job, const PrepInput& in,
std::error_code ec;
fs::create_directories(cand, ec);
char vf[64];
std::snprintf(vf, sizeof vf, "fps=%g", (double)job.video_fps * window);
std::snprintf(vf, sizeof vf, "fps=%g", (double)input_fps(job, in) * group);
// ffmpeg turns the picture by the container's matrix unless told not
// to, which is what the built-in decoder's auto_rotate matches.
std::vector<std::string> argv{job.ffmpeg_exe, "-nostdin", "-y"};
@@ -1456,11 +1464,19 @@ bool DatasetPrep::extract_video_ffmpeg(const PrepJob& job, const PrepInput& in,
if (window > 1) enter(Stage::Frames, lmsg::stage_select_sharpest.get());
fs::create_directories(out_dir, ec);
FrameSelectOptions so;
so.group = group;
so.max_frames = job.max_frames;
so.adaptive = job.adaptive_fps;
so.range = job.adaptive_range;
so.window = window;
if (fisheye) so.view = app::MotionView::Fisheye;
so.out_fov = fisheye ? 3.4034f : 1.5708f;
const int kept = select_sharpest_frames(
cand.string(), out_dir.string(), "", window, job.max_frames,
cand.string(), out_dir.string(), "", so,
[this](const std::string& l) { log(l); }, _cancel);
remove_tree(cand);
if (streams.size() > 1) fs::remove(track_path, ec);
if (streams > 1) fs::remove(track_path, ec);
if (kept < 0) {
error = _cancel.load() ? "cancelled" : "frame selection failed";
return false;
@@ -1556,6 +1572,9 @@ bool DatasetPrep::extract_360_ffmpeg(const PrepJob& job, const PrepInput& in,
const fs::path ws = job.workspace;
const int window = std::max(job.sharp_window, 1);
const int group = job.adaptive_fps
? std::max(window, (int)std::ceil(job.adaptive_range))
: window;
std::error_code ec;
// ffmpeg decodes both tracks and cuts the overlap strips out; the warp is
@@ -1567,7 +1586,7 @@ bool DatasetPrep::extract_360_ffmpeg(const PrepJob& job, const PrepInput& in,
remove_tree(cand);
fs::create_directories(cand, ec);
char pre[64];
std::snprintf(pre, sizeof pre, "fps=%g", (double)job.video_fps * window);
std::snprintf(pre, sizeof pre, "fps=%g", (double)input_fps(job, in) * group);
const std::string graph = app::pano360_graph(in.eac360, pre);
// A 360 capture's geometry is the EAC layout, not the display matrix: the
// built-in path leaves it alone and so must this one.
@@ -1587,8 +1606,17 @@ bool DatasetPrep::extract_360_ffmpeg(const PrepJob& job, const PrepInput& in,
const fs::path kept = ws / "canvas_tmp";
remove_tree(kept);
fs::create_directories(kept, ec);
FrameSelectOptions so;
so.group = group;
so.max_frames = job.max_frames;
so.adaptive = job.adaptive_fps;
so.range = job.adaptive_range;
so.window = window;
so.view = app::MotionView::Eac360;
so.eac = in.eac360;
so.out_fov = app::motion_out_fov(views);
const int n = select_sharpest_frames(
cand.string(), kept.string(), "", window, job.max_frames,
cand.string(), kept.string(), "", so,
[this](const std::string& l) { log(l); }, _cancel);
remove_tree(cand);
if (n < 0) {
+15 -1
View File
@@ -111,6 +111,10 @@ struct PrepInput {
// Which rig this input's lenses belong to (kRig*). A multi-lens video
// starts on its own; folders sharing a letter form one rig by file name.
int rig = kRigNone;
// Kept frames per second for THIS video; 0 takes the job's. A capture shot
// as several clips is rarely shot at one pace, and a clip walked through
// slowly wants fewer frames than the one that ran past the same wall.
float fps = 0.0f;
int video_tracks = 0; // 0 = not probed yet
// The camera folders found under this input, when it arrived with more
// than one. Empty means the lens above describes all of it.
@@ -210,7 +214,12 @@ struct PrepJob {
// panoramas and pinhole faces in one image tree describes no camera rig.
app::Pano360Options pano;
float video_fps = 2.0f; // kept frames per second
// Kept frames per second; PrepInput::fps overrides it per video.
float video_fps = 2.0f;
// Space them by view change rather than by time (app/FrameMotion.h): the
// rate above becomes the average and stays within `adaptive_range` of it.
bool adaptive_fps = false;
float adaptive_range = 4.0f;
int sharp_window = 3; // keep the sharpest of N (1 = off)
// Every track of a multi-lens file keeps the same instants (one sharpness
// window over all of them), so every frame is a rig frame. Built-in decoder only.
@@ -266,6 +275,11 @@ struct PrepJob {
std::string python_exe = "python3";
};
// The rate an input is actually extracted at: its own, or the job's.
inline float input_fps(const PrepJob& job, const PrepInput& in) {
return in.fps > 0.0f ? in.fps : job.video_fps;
}
// Images read where they are instead of gathered into the dataset's own
// images/ (see DatasetPrep::run). Several inputs reconstruct from ONE image
// tree, so there is nowhere for a second one to be read in place from.
+3
View File
@@ -23,6 +23,8 @@ namespace {
X("use_found_masks", use_found_masks) \
X("flip_found_masks", sfm.prep.flip_found_masks) \
X("video_fps", sfm.prep.video_fps) \
X("adaptive_fps", sfm.prep.adaptive_fps) \
X("adaptive_range", sfm.prep.adaptive_range) \
X("sharp_window", sfm.prep.sharp_window) \
X("sync_tracks", sfm.prep.sync_tracks) \
X("max_frames", sfm.prep.max_frames) \
@@ -197,6 +199,7 @@ void sanitize_dataset_settings(DatasetSettings& s) {
clamp_enum(p.pano.mode, 0, (int)app::Pano360Mode::Equirect);
clamp_to(p.pano.size, 0, 16384);
p.video_fps = std::clamp(p.video_fps, 0.0f, 240.0f);
p.adaptive_range = std::clamp(p.adaptive_range, 1.0f, 64.0f);
clamp_to(p.sharp_window, 1, 1000);
clamp_to(p.max_frames, 1, 1000000);
clamp_to(p.mask_detect_every, 1, 1000);
+166 -74
View File
@@ -12,6 +12,7 @@
#include <cstdint>
#include <cstdio>
#include <filesystem>
#include <memory>
#include <thread>
#include <vector>
@@ -21,33 +22,19 @@ namespace gui {
namespace {
// Sharpness metric, matching extract_frames.py add_frame(): resize to
// 512x512 (area average), grayscale, subtract mean, variance of the 3x3
// Laplacian. Returns -1 on decode failure.
double sharpness_score(const std::string& path) {
int W = 0, H = 0, C = 0;
std::vector<uint8_t> exr_rgb;
unsigned char* img = nullptr;
if (exr::is_exr(path)) {
exr::Info info;
if (!exr::decode_srgb8(path, exr::Options(), info, exr_rgb).empty()) return -1.0;
W = info.width;
H = info.height;
img = exr_rgb.data();
} else {
img = stbi_load(path.c_str(), &W, &H, &C, 3);
if (!img) return -1.0;
}
// Candidates loaded ahead of the motion tracker, which has to walk them in
// order. 64 grey frames is a few tens of megabytes whatever the source is.
constexpr size_t kChunk = 64;
constexpr int S = 512;
std::vector<float> gray((size_t)S * S);
double mean = 0.0;
for (int y = 0; y < S; y++) {
// Box-average the source rows/cols covered by this output cell.
int y0 = (int)((int64_t)y * H / S), y1 = (int)((int64_t)(y + 1) * H / S);
// Box-average an interleaved RGB image into a grey rectangle. `rows` is how
// much of the source to read, so an EAC canvas can be measured on its top row.
void box_grey(const unsigned char* img, int W, int rows, int ow, int oh,
float* out) {
for (int y = 0; y < oh; y++) {
int y0 = (int)((int64_t)y * rows / oh), y1 = (int)((int64_t)(y + 1) * rows / oh);
if (y1 <= y0) y1 = y0 + 1;
for (int x = 0; x < S; x++) {
int x0 = (int)((int64_t)x * W / S), x1 = (int)((int64_t)(x + 1) * W / S);
for (int x = 0; x < ow; x++) {
int x0 = (int)((int64_t)x * W / ow), x1 = (int)((int64_t)(x + 1) * W / ow);
if (x1 <= x0) x1 = x0 + 1;
double acc = 0.0;
for (int yy = y0; yy < y1; yy++)
@@ -56,36 +43,76 @@ double sharpness_score(const std::string& path) {
// BT.601 luma like cv2.cvtColor BGR2GRAY (RGB order here).
acc += 0.299 * p[0] + 0.587 * p[1] + 0.114 * p[2];
}
float v = (float)(acc / ((y1 - y0) * (x1 - x0)));
gray[(size_t)y * S + x] = v;
mean += v;
out[(size_t)y * ow + x] = (float)(acc / ((y1 - y0) * (x1 - x0)));
}
}
if (exr_rgb.empty()) stbi_image_free(img);
mean /= (double)S * S;
for (auto& v : gray) v -= (float)mean;
}
// Variance of the 3x3 Laplacian (cv2.Laplacian default kernel).
// Sharpness, matching extract_frames.py add_frame(): 512x512 grey, mean
// subtracted, variance of the 3x3 Laplacian.
double laplacian_variance(std::vector<float>& gray, int S) {
double mean = 0.0;
for (float v : gray) mean += v;
mean /= (double)gray.size();
for (float& v : gray) v -= (float)mean;
double sum = 0.0, sum2 = 0.0;
int64_t n = 0;
for (int y = 1; y < S - 1; y++)
for (int x = 1; x < S - 1; x++) {
const float* r = &gray[(size_t)y * S + x];
double lap = (double)r[-S] + r[S] + r[-1] + r[1] - 4.0 * r[0];
const double lap = (double)r[-S] + r[S] + r[-1] + r[1] - 4.0 * r[0];
sum += lap;
sum2 += lap * lap;
n++;
}
double mu = sum / n;
const double mu = sum / n;
return sum2 / n - mu * mu;
}
struct Analysis {
double score = -1.0;
std::vector<uint8_t> grey; // empty unless the motion pass wants it
};
// `mw`/`mh` <= 0 asks for the sharpness only. `top_rows` limits what the grey
// covers, for an EAC canvas whose bottom row is a different half of the sphere.
void analyze(const std::string& path, int mw, int mh, bool top_rows,
Analysis& out) {
int W = 0, H = 0, C = 0;
std::vector<uint8_t> exr_rgb;
unsigned char* img = nullptr;
if (exr::is_exr(path)) {
exr::Info info;
if (!exr::decode_srgb8(path, exr::Options(), info, exr_rgb).empty()) return;
W = info.width;
H = info.height;
img = exr_rgb.data();
} else {
img = stbi_load(path.c_str(), &W, &H, &C, 3);
if (!img) return;
}
constexpr int S = 512;
std::vector<float> gray((size_t)S * S);
box_grey(img, W, H, S, S, gray.data());
out.score = laplacian_variance(gray, S);
if (mw > 0 && mh > 0) {
std::vector<float> f((size_t)mw * mh);
box_grey(img, W, top_rows ? H / 2 : H, mw, mh, f.data());
out.grey.resize(f.size());
for (size_t i = 0; i < f.size(); i++)
out.grey[i] = (uint8_t)std::min(255.0f, std::max(0.0f, f[i] + 0.5f));
}
if (exr_rgb.empty()) stbi_image_free(img);
}
} // namespace
int select_sharpest_frames(const std::string& cand_dir,
const std::string& out_dir,
const std::string& prefix,
int group, int max_frames,
const FrameSelectOptions& options,
const std::function<void(const std::string&)>& log,
const std::atomic<bool>& cancel) {
std::error_code ec;
@@ -95,49 +122,113 @@ int select_sharpest_frames(const std::string& cand_dir,
std::sort(files.begin(), files.end());
if (files.empty()) return -1;
group = std::max(group, 1);
if (max_frames > 0)
group = std::max<int>(group, (int)((files.size() + max_frames - 1) / max_frames));
int group = std::max(options.group, 1);
if (!options.adaptive && options.max_frames > 0)
group = std::max<int>(
group, (int)((files.size() + options.max_frames - 1) / options.max_frames));
// Score all candidates in parallel.
std::vector<double> scores(files.size(), -1.0);
if (group > 1) {
unsigned n_threads = std::max(1u, std::thread::hardware_concurrency());
std::atomic<size_t> next{0};
std::atomic<size_t> done{0};
std::vector<std::thread> pool;
for (unsigned t = 0; t < n_threads; t++)
pool.emplace_back([&] {
for (;;) {
size_t i = next.fetch_add(1);
if (i >= files.size() || cancel.load()) return;
// A decode that throws would be std::terminate off a
// worker thread. An unscored frame just loses its group.
try {
scores[i] = sharpness_score(files[i].string());
} catch (...) {}
done.fetch_add(1);
}
});
// Progress line roughly once a second.
while (done.load() < files.size() && !cancel.load()) {
std::this_thread::sleep_for(std::chrono::milliseconds(1000));
if (log) log("scored " + std::to_string(done.load()) + "/" +
std::to_string(files.size()) + " candidate frames");
// The motion tracker's frames: the EAC top row is three faces wide and one
// tall, with no overlap strips left in a canvas ffmpeg already cut.
app::MotionOptions mo;
std::unique_ptr<app::MotionTracker> tracker;
if (options.adaptive) {
mo.view = options.view;
mo.out_fov = options.out_fov;
int src_w = 0, src_h = 0;
if (options.eac.valid()) {
mo.eac.face = options.eac.face;
mo.eac.strip = 0;
mo.eac.track_w = 3 * options.eac.face;
mo.eac.track_h = options.eac.face;
src_w = mo.eac.track_w;
src_h = mo.eac.track_h;
} else {
int w = 0, h = 0, c = 0;
if (stbi_info(files[0].string().c_str(), &w, &h, &c)) {
src_w = w;
src_h = h;
}
}
for (auto& th : pool) th.join();
if (cancel.load()) return -1;
app::motion_frame_size(mo.view, src_w, src_h, mo.width, mo.height);
if (mo.width > 0) tracker = std::make_unique<app::MotionTracker>(mo);
}
// Keep the sharpest per group; move to out_dir, delete the rest.
// Score, and measure, in chunks: the tracker has to see the frames in
// order, and everything before it parallelizes.
std::vector<double> scores(files.size(), -1.0);
const bool want_scores = group > 1 || options.adaptive;
if (want_scores) {
const unsigned n_threads = std::max(1u, std::thread::hardware_concurrency());
auto last_log = std::chrono::steady_clock::now();
for (size_t base = 0; base < files.size(); base += kChunk) {
if (cancel.load()) return -1;
const size_t end = std::min(base + kChunk, files.size());
std::vector<Analysis> got(end - base);
std::atomic<size_t> next{base};
std::vector<std::thread> pool;
for (unsigned t = 0; t < n_threads; t++)
pool.emplace_back([&] {
for (;;) {
const size_t i = next.fetch_add(1);
if (i >= end || cancel.load()) return;
// A decode that throws would be std::terminate off a
// worker thread. An unscored frame loses its group.
try {
analyze(files[i].string(), tracker ? mo.width : 0,
tracker ? mo.height : 0, options.eac.valid(),
got[i - base]);
} catch (...) {}
}
});
for (auto& th : pool) th.join();
if (cancel.load()) return -1;
for (size_t i = base; i < end; i++) {
scores[i] = got[i - base].score;
if (tracker && !got[i - base].grey.empty())
tracker->track(got[i - base].grey.data(), (int64_t)i);
}
const auto now = std::chrono::steady_clock::now();
if (log && (end == files.size() ||
now - last_log > std::chrono::seconds(1))) {
last_log = now;
log("scored " + std::to_string(end) + "/" +
std::to_string(files.size()) + " candidate frames");
}
}
}
// Which candidates survive: one per group, or one per equal share of view
// change with the sharpest of the window around it.
std::vector<size_t> keep;
if (tracker) {
tracker->finish();
const int window = std::max(options.window, 1);
const std::vector<int64_t> plan = app::plan_by_motion(
tracker->costs(), tracker->ends(), (int64_t)files.size(), group,
window, options.range, options.max_frames);
for (int64_t at : plan) {
const size_t g1 = (size_t)at + 1;
const size_t g0 = g1 > (size_t)window ? g1 - (size_t)window : 0;
size_t best = g0;
for (size_t i = g0 + 1; i < g1; i++)
if (scores[i] > scores[best]) best = i;
if (keep.empty() || keep.back() != best) keep.push_back(best);
}
}
if (keep.empty())
for (size_t g0 = 0; g0 < files.size(); g0 += (size_t)group) {
const size_t g1 = std::min(g0 + (size_t)group, files.size());
size_t best = g0;
for (size_t i = g0 + 1; i < g1; i++)
if (scores[i] > scores[best]) best = i;
keep.push_back(best);
}
fs::create_directories(out_dir, ec);
std::vector<bool> kept_flag(files.size(), false);
int kept = 0;
for (size_t g0 = 0; g0 < files.size(); g0 += group) {
size_t g1 = std::min(g0 + (size_t)group, files.size());
size_t best = g0;
for (size_t i = g0 + 1; i < g1; i++)
if (scores[i] > scores[best]) best = i;
std::string ext = files[best].extension().string();
for (size_t best : keep) {
const std::string ext = files[best].extension().string();
char name[64];
std::snprintf(name, sizeof name, "%s%05d%s", prefix.c_str(), kept,
ext.empty() ? ".jpg" : ext.c_str());
@@ -147,11 +238,12 @@ int select_sharpest_frames(const std::string& cand_dir,
fs::copy_options::overwrite_existing, ec);
fs::remove(files[best], ec);
}
kept_flag[best] = true;
kept++;
for (size_t i = g0; i < g1; i++)
if (i != best) fs::remove(files[i], ec);
if (cancel.load()) return -1;
}
for (size_t i = 0; i < files.size(); i++)
if (!kept_flag[i]) fs::remove(files[i], ec);
return kept;
}
+26 -15
View File
@@ -1,15 +1,14 @@
#pragma once
// FrameSelect -- motion-blur-aware frame selection for video datasets.
// C++ multithreaded port of reference/scripts/extract_frames.py's
// FrameSelector: the sharpness metric is the variance of the 3x3 Laplacian
// of the mean-subtracted 512x512 grayscale image; within each group of
// `group` consecutive candidate frames the sharpest one is kept.
// FrameSelect -- which of the candidate frames ffmpeg extracted are kept.
//
// The GUI extracts candidates with ffmpeg at (target fps x group), then
// calls this to keep the best frame per group -- same output rate as a
// plain fps filter, but each kept frame is the least blurry of its
// neighborhood.
// The sharpness metric is the variance of the 3x3 Laplacian of the
// mean-subtracted 512x512 grayscale image. Evenly, one per group of `group`
// candidates; or spaced by how much the view changes (app/FrameMotion.h),
// which is what the built-in decoder's adaptive mode does from the source
// frames themselves.
#include "app/FrameMotion.h"
#include <atomic>
#include <functional>
@@ -17,16 +16,28 @@
namespace gui {
struct FrameSelectOptions {
int group = 1; // candidates per kept frame
int max_frames = 0; // 0 = no cap
// Adaptive: the rate stays within `range` either side of the average, and
// the sharpest of `window` candidates around each chosen one is kept.
bool adaptive = false;
float range = 4.0f;
int window = 3;
app::MotionView view = app::MotionView::Planar;
// Set when the candidates are EAC canvases: only the top row is measured,
// which is half the sphere and all a rotation needs.
app::Eac360Layout eac;
float out_fov = 1.5708f;
};
// Scores every image in cand_dir (sorted by filename) across all hardware
// threads, keeps the sharpest of each consecutive `group`, and MOVES the
// keepers to out_dir as <prefix>NNNNN.<ext> (contiguous numbering). Losers
// are deleted. If keeping one per group still exceeds max_frames, the group
// size is enlarged. Returns the number of frames kept, or -1 on error /
// cancellation.
// threads and MOVES the keepers to out_dir as <prefix>NNNNN.<ext>, contiguously
// numbered; the losers are deleted. -1 on error or cancellation.
int select_sharpest_frames(const std::string& cand_dir,
const std::string& out_dir,
const std::string& prefix,
int group, int max_frames,
const FrameSelectOptions& options,
const std::function<void(const std::string&)>& log,
const std::atomic<bool>& cancel);
+47 -2
View File
@@ -520,6 +520,12 @@ void GuiApp::write_run_settings(std::ofstream& f) {
line("force_external_decode", cfg_str(j.prep.force_external_decode));
line("force_external_masking", cfg_str(j.prep.force_external_masking));
line("video_fps", cfg_str(j.prep.video_fps));
line("adaptive_fps", cfg_str(j.prep.adaptive_fps));
if (j.prep.adaptive_fps) line("adaptive_range", cfg_str(j.prep.adaptive_range));
for (const PrepInput& s : _sources)
if (s.is_video && s.fps > 0.0f)
line("video_fps:" + (s.subdir.empty() ? s.path : s.subdir),
cfg_str(s.fps));
line("sharp_window", std::to_string(j.prep.sharp_window));
line("sync_tracks", cfg_str(j.prep.sync_tracks));
line("max_frames", std::to_string(j.prep.max_frames));
@@ -2690,6 +2696,8 @@ void GuiApp::sync_dataset_jobs() {
prep.workspace = _workspace;
prep.resume = _resume;
prep.video_fps = _sfm_job.prep.video_fps;
prep.adaptive_fps = _sfm_job.prep.adaptive_fps;
prep.adaptive_range = _sfm_job.prep.adaptive_range;
prep.sharp_window = _sfm_job.prep.sharp_window;
prep.pano = _sfm_job.prep.pano;
prep.max_frames = _sfm_job.prep.max_frames;
@@ -2721,6 +2729,8 @@ void GuiApp::sync_dataset_jobs() {
_colmap_job.workspace = prep.workspace;
_colmap_job.resume = prep.resume;
_colmap_job.video_fps = prep.video_fps;
_colmap_job.adaptive_fps = prep.adaptive_fps;
_colmap_job.adaptive_range = prep.adaptive_range;
_colmap_job.sharp_window = prep.sharp_window;
_colmap_job.pano = prep.pano;
_colmap_job.max_frames = prep.max_frames;
@@ -2906,12 +2916,17 @@ void GuiApp::draw_dataset_source() {
// its own folder of frames under images/, and so its own camera.
int remove = -1;
bool edited = false;
// Several clips are rarely shot at one pace, so each gets its own rate
// once there is another one to differ from.
int n_video = 0;
for (const PrepInput& s : _sources) n_video += s.is_video ? 1 : 0;
const bool per_video_fps = n_video > 1;
for (size_t i = 0; i < _sources.size(); i++) {
PrepInput& s = _sources[i];
ImGui::PushID((int)i);
// Room for Browse + Remove + where the frames go + what sensors it
// carries, which is longer than any other row on the screen.
ImGui::SetNextItemWidth(px(-500.0f));
ImGui::SetNextItemWidth(px(per_video_fps ? -580.0f : -500.0f));
if (ui::InputTextRaw("##in", &s.path)) {
std::error_code ec;
s.is_video = !fs::is_directory(s.path, ec) && is_video_path(s.path);
@@ -2937,6 +2952,24 @@ void GuiApp::draw_dataset_source() {
}
ImGui::SameLine();
if (ui::Button(dmsg::remove)) remove = (int)i;
if (per_video_fps) {
ImGui::SameLine();
if (s.is_video) {
ImGui::BeginDisabled(dataset_locked(Stage::Frames));
ImGui::SetNextItemWidth(px(74.0f));
float shown = input_fps(_sfm_job.prep, s);
// Typing the dataset-wide rate back in is how a row goes back
// to following it, which is what an empty override means.
if (ui::InputFloatRaw("##fps", &shown, "%.3g fps")) {
if (shown <= 0.0f) shown = _sfm_job.prep.video_fps;
s.fps = shown == _sfm_job.prep.video_fps ? 0.0f : shown;
}
ui::help_on_hover(dmsg::video_fps_this_one_help);
ImGui::EndDisabled();
} else {
ImGui::Dummy(ImVec2(px(74.0f), 0.0f));
}
}
ImGui::SameLine();
// What this input is, and -- the part worth seeing before pressing the
// button -- whether masks were found for it. Four whole messages
@@ -3245,7 +3278,19 @@ void GuiApp::draw_dataset_basics() {
ImGui::SetNextItemWidth(px(220.0f));
ui::InputFloat(dmsg::frames_per_second, &_sfm_job.prep.video_fps,
0, 0, "%.2g");
ui::help_on_hover(dmsg::frames_per_second_help);
ui::help_on_hover(_sfm_job.prep.adaptive_fps
? dmsg::frames_per_second_help_adaptive
: dmsg::frames_per_second_help);
ui::Checkbox(dmsg::adaptive_fps, &_sfm_job.prep.adaptive_fps);
ui::help_on_hover(dmsg::adaptive_fps_help);
if (_sfm_job.prep.adaptive_fps) {
ImGui::Indent();
ImGui::SetNextItemWidth(px(220.0f));
ui::SliderFloat(dmsg::adaptive_range, &_sfm_job.prep.adaptive_range,
1.0f, 16.0f, "%.1f");
ui::help_on_hover(dmsg::adaptive_range_help);
ImGui::Unindent();
}
ImGui::SetNextItemWidth(px(220.0f));
ui::SliderInt(dmsg::sharpness_window, &_sfm_job.prep.sharp_window, 1, 8);
ui::help_on_hover(dmsg::sharpness_window_help);
+6 -1
View File
@@ -388,7 +388,12 @@ void dataset_adapt_preset(const std::string& preset,
void resolve_source_lenses(std::vector<PrepInput>& sources, SfmJob& sfm,
ColmapJob& colmap) {
if (sources.empty()) return;
if (any_pano360(sources) && sfm.prep.pano.mode != app::Pano360Mode::Off) {
if (any_pano360(sources)) {
// Nothing on the panel can say "leave the packed tracks alone", and
// the combo draws that state as "views", so a preset saved off a flat
// capture must not be able to leave a 360 one in it.
if (sfm.prep.pano.mode == app::Pano360Mode::Off)
sfm.prep.pano.mode = app::Pano360Mode::Faces;
if (sfm.prep.pano.size <= 0) reset_pano_size(sources, sfm.prep.pano);
apply_pano_lens(sources, sfm, colmap);
return;
@@ -57,6 +57,8 @@ static void test_dataset_preset() {
s.sfm.prep.photo_import = gui::PhotoImport::Move;
s.sfm.prep.flip_found_masks = true;
s.sfm.prep.video_fps = 5.5f;
s.sfm.prep.adaptive_fps = true;
s.sfm.prep.adaptive_range = 2.5f;
s.sfm.prep.sharp_window = 7;
s.sfm.prep.sync_tracks = false;
s.sfm.prep.max_frames = 1234;
@@ -167,6 +169,8 @@ static void test_dataset_preset() {
CHECK(b.sfm.prep.photo_import == s.sfm.prep.photo_import);
CHECK_EQ(b.sfm.prep.flip_found_masks, s.sfm.prep.flip_found_masks);
CHECK_EQ(b.sfm.prep.video_fps, s.sfm.prep.video_fps);
CHECK_EQ(b.sfm.prep.adaptive_fps, s.sfm.prep.adaptive_fps);
CHECK_EQ(b.sfm.prep.adaptive_range, s.sfm.prep.adaptive_range);
CHECK_EQ(b.sfm.prep.sharp_window, s.sfm.prep.sharp_window);
CHECK_EQ(b.sfm.prep.sync_tracks, s.sfm.prep.sync_tracks);
CHECK_EQ(b.sfm.prep.max_frames, s.sfm.prep.max_frames);
@@ -323,6 +327,7 @@ static void test_sanitize() {
s.sfm.features = 0;
s.sfm.matcher = 1; // LightGlue without a learned frontend
s.sfm.prep.sharp_window = -3;
s.sfm.prep.adaptive_range = 0.1f;
s.mask.threshold = 4.0f;
s.colmap.matcher = 0;
s.colmap.camera_model = "NONSENSE";
@@ -331,6 +336,7 @@ static void test_sanitize() {
CHECK_EQ(s.sfm.camera_model, std::string("opencv"));
CHECK_EQ(s.sfm.matcher, 0);
CHECK(s.sfm.prep.sharp_window >= 1);
CHECK(s.sfm.prep.adaptive_range >= 1.0f);
CHECK(s.mask.threshold <= 1.0f);
CHECK(s.colmap.matcher >= 1);
CHECK_EQ(s.colmap.camera_model, std::string("OPENCV"));
+119
View File
@@ -0,0 +1,119 @@
// frame_motion -- the adaptive frame plan (app/FrameMotion.h). Three things
// have gone wrong here and each is silent, so each is asserted: the count has
// to land on the budget, the gaps have to stay inside the rate bounds, and a
// burst of motion must not swallow the budget it cannot spend.
#include "app/FrameMotion.h"
#include <cmath>
#include <cstdio>
#include <string>
#include <vector>
namespace {
int g_failures = 0;
void expect(bool ok, const std::string& what) {
std::printf("%s %s\n", ok ? "ok " : "BAD ", what.c_str());
if (!ok) g_failures++;
}
// One sample per source frame, so the plan is in the units the test states.
std::vector<int64_t> indices(size_t n) {
std::vector<int64_t> v(n);
for (size_t i = 0; i < n; i++) v[i] = (int64_t)i;
return v;
}
void check_plan(const char* name, const std::vector<float>& cost, int skip,
int window, float range, int64_t want) {
const std::vector<int64_t> at = indices(cost.size());
const std::vector<int64_t> plan = app::plan_by_motion(
cost, at, (int64_t)cost.size(), skip, window, range, 0);
const std::string tag(name);
// Bisection lands on the budget or just under it; a tenth either way is
// the granularity of a plan that can only place frames on samples.
expect((int64_t)plan.size() <= want &&
(int64_t)plan.size() >= want - want / 10 - 1,
tag + ": " + std::to_string(plan.size()) + " frames for a budget of " +
std::to_string(want));
const int64_t min_gap = std::max<int64_t>(window, (int64_t)(skip / range));
const int64_t max_gap = std::max<int64_t>(min_gap, (int64_t)(skip * range));
int64_t tight = (int64_t)cost.size(), wide = 0;
for (size_t i = 1; i < plan.size(); i++) {
const int64_t gap = plan[i] - plan[i - 1];
tight = std::min(tight, gap);
wide = std::max(wide, gap);
}
expect(plan.size() < 2 || (tight >= min_gap && wide <= max_gap),
tag + ": gaps " + std::to_string(tight) + ".." + std::to_string(wide) +
" within " + std::to_string(min_gap) + ".." + std::to_string(max_gap));
}
} // namespace
int main() {
// Even motion: the plan should be the fixed schedule in all but name.
{
const std::vector<float> cost(600, 0.03f);
check_plan("even", cost, 15, 3, 4.0f, 40);
const std::vector<int64_t> plan = app::plan_by_motion(
cost, indices(cost.size()), 600, 15, 3, 4.0f, 0);
int64_t wide = 0, tight = 600;
for (size_t i = 1; i < plan.size(); i++) {
wide = std::max(wide, plan[i] - plan[i - 1]);
tight = std::min(tight, plan[i] - plan[i - 1]);
}
expect(wide - tight <= 1, "even: the spacing stays even");
}
// Half the clip still, half moving: the moving half takes the frames.
{
std::vector<float> cost(600, 0.002f);
for (size_t i = 300; i < 600; i++) cost[i] = 0.08f;
check_plan("split", cost, 15, 3, 4.0f, 40);
const std::vector<int64_t> plan = app::plan_by_motion(
cost, indices(cost.size()), 600, 15, 3, 4.0f, 0);
int first = 0;
for (int64_t at : plan) first += at < 300 ? 1 : 0;
expect(first * 2 < (int)plan.size() - first,
"split: the moving half gets more than twice the frames");
}
// A burst one sample wide: its budget cannot be spent inside it, and the
// rest of the clip must not be starved for it.
{
std::vector<float> cost(600, 0.01f);
cost[200] = 40.0f;
check_plan("burst", cost, 15, 3, 4.0f, 40);
}
// A tripod: nothing changes, so nothing justifies more than the slowest
// rate the bounds allow.
{
const std::vector<float> cost(600, 0.0f);
const std::vector<int64_t> plan = app::plan_by_motion(
cost, indices(cost.size()), 600, 15, 3, 4.0f, 0);
expect((int)plan.size() <= 600 / 60 + 1,
"still: " + std::to_string(plan.size()) +
" frames, at most the slowest rate");
}
// The cap is a cap, not a target.
{
const std::vector<float> cost(600, 0.03f);
const std::vector<int64_t> plan = app::plan_by_motion(
cost, indices(cost.size()), 600, 15, 3, 4.0f, 12);
expect((int)plan.size() <= 12, "cap: " + std::to_string(plan.size()) +
" frames within a cap of 12");
expect((int)plan.size() >= 10, "cap: the cap is nearly filled");
}
// Nothing to plan from.
expect(app::plan_by_motion({}, {}, 0, 15, 3, 4.0f, 0).empty(),
"empty: no samples, no plan");
std::printf("%s\n", g_failures ? "FAILED" : "PASSED");
return g_failures ? 1 : 0;
}
+30
View File
@@ -901,6 +901,36 @@ SS_MSG(xs_measured,
RU("измерено кадров"),
TR("ölçülen kare"));
SS_MSG(xs_analyzed,
EN("frames analyzed for motion"),
JA("動きを解析したフレーム"),
ZH_HANS("已分析运动的帧"),
ZH_HANT("已分析運動的影格"),
KO("움직임을 분석한 프레임"),
DE("auf Bewegung geprüfte Einzelbilder"),
FR("images analysées pour le mouvement"),
ES("fotogramas analizados en movimiento"),
PT("quadros analisados quanto ao movimento"),
IT("fotogrammi analizzati per il movimento"),
NL("op beweging geanalyseerde beelden"),
RU("кадров проанализировано на движение"),
TR("hareket için incelenen kare"));
SS_MSG(xs_motion,
EN("motion pass"),
JA("動き解析パス"),
ZH_HANS("运动分析遍"),
ZH_HANT("運動分析階段"),
KO("움직임 분석 패스"),
DE("Bewegungsdurchlauf"),
FR("passe de mouvement"),
ES("pasada de movimiento"),
PT("passagem de movimento"),
IT("passaggio di movimento"),
NL("bewegingsdoorloop"),
RU("проход по движению"),
TR("hareket geçişi"));
SS_MSG(xs_written,
EN("frames written"),
JA("書き出したフレーム"),
+177 -17
View File
@@ -2520,42 +2520,202 @@ SS_MSG(frames_per_second,
SS_MSG(frames_per_second_help,
EN("How many frames to keep per second of video. 1-3 is right for a slow "
"walkthrough; more only helps if the camera moved fast. Applies to "
"every video in the list."),
"walkthrough; more only helps if the camera moved fast. A video with a "
"rate of its own uses that instead."),
JA("動画1秒あたり何フレーム残すかです。ゆっくり歩いて撮ったなら 1〜3 が"
"適切で、それ以上が効くのはカメラが速く動いたときだけです。リスト内の"
"すべての動画に適用されます。"),
"適切で、それ以上が効くのはカメラが速く動いたときだけです。個別の値を"
"入れた動画はそちらに従います。"),
ZH_HANS("每秒视频保留多少帧。慢慢走着拍的话 1-3 就合适;更高只有在相机移动"
"很快时才有用。对列表中的所有视频都生效。"),
"很快时才有用。单独设了帧率的视频按各自的来。"),
ZH_HANT("每秒影片保留多少影格。慢慢走著拍的話 1-3 就合適;更高只有在相機移動"
"很快時才有用。對清單中的所有影片都生效。"),
"很快時才有用。單獨設了影格率的影片按各自的來。"),
KO("동영상 1초당 몇 프레임을 남길지입니다. 천천히 걸으며 찍었다면 1~3이 "
"알맞고, 그보다 높이는 건 카메라가 빠르게 움직였을 때만 도움이 됩니다. "
"목록의 모든 동영상에 적용됩니다."),
"자체 값이 있는 동영상은 그 값을 씁니다."),
DE("Wie viele Bilder je Sekunde Video behalten werden. 1-3 passt für "
"einen langsamen Rundgang; mehr hilft nur, wenn die Kamera schnell "
"bewegt wurde. Gilt für jedes Video in der Liste."),
"bewegt wurde. Ein Video mit eigener Rate nimmt seine eigene."),
FR("Combien d'images conserver par seconde de vidéo. 1 à 3 convient à une "
"déambulation lente ; davantage n'aide que si la caméra bougeait vite. "
"S'applique à toutes les vidéos de la liste."),
"Une vidéo ayant son propre débit garde le sien."),
ES("Cuántos fotogramas conservar por segundo de vídeo. De 1 a 3 va bien "
"para un recorrido lento; más solo ayuda si la cámara se movía rápido. "
"Se aplica a todos los vídeos de la lista."),
"Un vídeo con su propia tasa usa la suya."),
PT("Quantos quadros manter por segundo de vídeo. De 1 a 3 serve para um "
"percurso lento; mais só ajuda se a câmera se moveu rápido. Vale para "
"todos os vídeos da lista."),
"percurso lento; mais só ajuda se a câmera se moveu rápido. Um vídeo "
"com taxa própria usa a dele."),
IT("Quanti fotogrammi tenere per ogni secondo di video. Da 1 a 3 va bene "
"per una camminata lenta; di più serve solo se la fotocamera si "
"muoveva in fretta. Vale per tutti i video dell'elenco."),
"muoveva in fretta. Un video con una frequenza propria usa la sua."),
NL("Hoeveel beelden per seconde video bewaard blijven. 1-3 past bij een "
"rustige rondgang; meer helpt alleen als de camera snel bewoog. Geldt "
"voor elke video in de lijst."),
"rustige rondgang; meer helpt alleen als de camera snel bewoog. Een "
"video met een eigen tempo houdt dat van zichzelf."),
RU("Сколько кадров оставлять на секунду видео. 1-3 подходит для "
"неторопливого обхода; больше помогает, только если камера двигалась "
"быстро. Применяется ко всем видео в списке."),
"быстро. Видео со своей частотой берёт свою."),
TR("Videonun her saniyesinden kaç karenin tutulacağı. Yavaş bir gezinti "
"için 1-3 uygundur; daha fazlası yalnızca kamera hızlı hareket ettiyse "
"işe yarar. Listedeki bütün videolara uygulanır."));
"işe yarar. Kendi hızı olan video kendininkini kullanır."));
SS_MSG(frames_per_second_help_adaptive,
EN("The AVERAGE number of frames to keep per second of video; where they "
"fall is decided by how much the view changes. A video with a rate of "
"its own uses that instead."),
JA("動画1秒あたり平均で何フレーム残すかです。どこで残すかは見えの変化量が"
"決めます。個別の値を入れた動画はそちらに従います。"),
ZH_HANS("每秒视频平均保留多少帧;具体取在哪里由画面变化量决定。单独设了帧率"
"的视频按各自的来。"),
ZH_HANT("每秒影片平均保留多少影格;具體取在哪裡由畫面變化量決定。單獨設了影格率"
"的影片按各自的來。"),
KO("동영상 1초당 평균 몇 프레임을 남길지입니다. 어디서 남길지는 시야가 바뀐 "
"정도가 정합니다. 자체 값이 있는 동영상은 그 값을 씁니다."),
DE("Wie viele Bilder je Sekunde Video im DURCHSCHNITT behalten werden; wo "
"sie liegen, entscheidet die Änderung des Blicks. Ein Video mit eigener "
"Rate nimmt seine eigene."),
FR("Le nombre MOYEN d'images conservées par seconde de vidéo ; leur "
"emplacement suit le changement de vue. Une vidéo ayant son propre "
"débit garde le sien."),
ES("El número MEDIO de fotogramas conservados por segundo de vídeo; dónde "
"caen lo decide cuánto cambia la vista. Un vídeo con su propia tasa usa "
"la suya."),
PT("O número MÉDIO de quadros guardados por segundo de vídeo; onde caem "
"depende de quanto a vista muda. Um vídeo com taxa própria usa o dele."),
IT("Il numero MEDIO di fotogrammi tenuti per secondo di video; dove "
"cadono lo decide quanto cambia la vista. Un video con una frequenza "
"propria usa la sua."),
NL("Het GEMIDDELDE aantal beelden per seconde video; waar ze vallen "
"bepaalt hoeveel het beeld verandert. Een video met een eigen tempo "
"houdt dat van zichzelf."),
RU("СРЕДНЕЕ число кадров, оставляемых на секунду видео; где именно они "
"придутся, решает изменение вида. Видео со своей частотой берёт свою."),
TR("Videonun her saniyesinden ORTALAMA kaç kare tutulacağı; nereye "
"düşecekleri görüntünün ne kadar değiştiğine bağlıdır. Kendi hızı olan "
"video kendininkini kullanır."));
SS_MSG(video_fps_this_one_help,
EN("Frames per second for this video alone. Set it back to the rate above "
"to follow that one again."),
JA("この動画だけの毎秒フレーム数です。上と同じ値に戻すと、また上に従います。"),
ZH_HANS("仅用于这个视频的每秒帧数。改回上面的值就重新跟随上面的设置。"),
ZH_HANT("僅用於這個影片的每秒影格數。改回上面的值就重新跟隨上面的設定。"),
KO("이 동영상에만 적용되는 초당 프레임 수입니다. 위의 값으로 되돌리면 다시 "
"위를 따릅니다."),
DE("Bilder je Sekunde nur für dieses Video. Auf die Rate oben zurückgesetzt "
"folgt es wieder jener."),
FR("Images par seconde pour cette vidéo seule. Remettez le débit ci-dessus "
"pour qu'elle le suive à nouveau."),
ES("Fotogramas por segundo solo para este vídeo. Vuelve a poner la tasa de "
"arriba para que la siga otra vez."),
PT("Quadros por segundo só para este vídeo. Reponha a taxa acima para que "
"volte a segui-la."),
IT("Fotogrammi al secondo solo per questo video. Rimetti la frequenza "
"qui sopra perché la segua di nuovo."),
NL("Beelden per seconde alleen voor deze video. Zet het terug op het tempo "
"hierboven om dat weer te volgen."),
RU("Кадров в секунду только для этого видео. Верните значение сверху, "
"чтобы снова следовать ему."),
TR("Yalnızca bu video için saniyedeki kare sayısı. Yukarıdaki hıza geri "
"ayarlayınca yine onu izler."));
SS_MSG(adaptive_fps,
EN("Adapt the rate to the motion"),
JA("動きに合わせてレートを変える"),
ZH_HANS("按运动调整帧率"),
ZH_HANT("依運動調整影格率"),
KO("움직임에 맞춰 속도 조절"),
DE("Rate an die Bewegung anpassen"),
FR("Adapter le débit au mouvement"),
ES("Adaptar la tasa al movimiento"),
PT("Adaptar a taxa ao movimento"),
IT("Adatta la frequenza al movimento"),
NL("Tempo aanpassen aan de beweging"),
RU("Подстраивать частоту под движение"),
TR("Hızı harekete göre ayarla"));
SS_MSG(adaptive_fps_help,
EN("Keep more frames where the camera moves fast or passes close to "
"something, fewer where it only turns on the spot or looks at distant "
"scenery. The rate above becomes the average. Costs one extra pass "
"over each video."),
JA("カメラが速く動いたときや近くの物のそばを通ったときは多めに、その場で"
"向きを変えただけのときや遠景を見ているときは少なめに残します。上の"
"レートは平均値になります。動画ごとに1回分の解析が余計にかかります。"),
ZH_HANS("相机移动快或贴近物体时多留几帧,原地转动或只看远景时少留。上面的帧率"
"变成平均值。每个视频要多跑一遍分析。"),
ZH_HANT("相機移動快或貼近物體時多留幾格,原地轉動或只看遠景時少留。上面的影格率"
"變成平均值。每個影片要多跑一遍分析。"),
KO("카메라가 빠르게 움직이거나 가까운 물체를 지날 때는 더 많이, 제자리에서 "
"돌거나 먼 풍경만 볼 때는 더 적게 남깁니다. 위의 속도는 평균이 됩니다. "
"동영상마다 분석 패스가 한 번 더 듭니다."),
DE("Mehr Bilder behalten, wo die Kamera schnell fährt oder dicht an etwas "
"vorbeikommt, weniger, wo sie sich nur dreht oder in die Ferne sieht. "
"Die Rate oben wird der Durchschnitt. Kostet einen zusätzlichen "
"Durchlauf je Video."),
FR("Conserver davantage d'images là où la caméra va vite ou frôle un "
"objet, moins là où elle pivote sur place ou regarde au loin. Le débit "
"ci-dessus devient la moyenne. Coûte une passe supplémentaire par "
"vidéo."),
ES("Conservar más fotogramas donde la cámara va rápido o pasa cerca de "
"algo, y menos donde solo gira sobre sí misma o mira a lo lejos. La "
"tasa de arriba pasa a ser el promedio. Cuesta una pasada más por "
"vídeo."),
PT("Guardar mais quadros onde a câmera anda depressa ou passa perto de "
"algo, e menos onde apenas gira no lugar ou olha ao longe. A taxa "
"acima passa a ser a média. Custa uma passagem extra por vídeo."),
IT("Tenere più fotogrammi dove la camera va veloce o sfiora qualcosa, "
"meno dove ruota sul posto o guarda lontano. La frequenza qui sopra "
"diventa la media. Costa un passaggio in più per video."),
NL("Meer beelden bewaren waar de camera snel gaat of vlak langs iets "
"komt, minder waar hij alleen draait of in de verte kijkt. Het tempo "
"hierboven wordt het gemiddelde. Kost één extra doorloop per video."),
RU("Оставлять больше кадров там, где камера идёт быстро или проходит "
"близко к предмету, и меньше там, где она лишь поворачивается на месте "
"или смотрит вдаль. Частота сверху становится средней. Стоит одного "
"дополнительного прохода на каждое видео."),
TR("Kamera hızlı giderken ya da bir şeyin yakınından geçerken daha çok, "
"yerinde dönerken ya da uzağa bakarken daha az kare tut. Yukarıdaki "
"hız ortalama olur. Her video için bir ek geçişe mal olur."));
SS_MSG(adaptive_range,
EN("Spread"),
JA("振れ幅"),
ZH_HANS("浮动范围"),
ZH_HANT("浮動範圍"),
KO("변동 폭"),
DE("Spanne"),
FR("Amplitude"),
ES("Margen"),
PT("Margem"),
IT("Escursione"),
NL("Spreiding"),
RU("Разброс"),
TR("Aralık"));
SS_MSG(adaptive_range_help,
EN("How far the rate may stray from the average, either way. 4 lets it "
"run between a quarter of it and four times it."),
JA("レートが平均からどこまで離れてよいかです。4 なら平均の 1/4 から 4 倍まで"
"振れます。"),
ZH_HANS("帧率相对平均值的上下浮动倍数。设为 4 表示可在平均值的 1/4 到 4 倍之间。"),
ZH_HANT("影格率相對平均值的上下浮動倍數。設為 4 表示可在平均值的 1/4 到 4 倍之間。"),
KO("속도가 평균에서 얼마나 벗어날 수 있는지입니다. 4면 평균의 1/4에서 4배 "
"사이를 오갑니다."),
DE("Wie weit die Rate nach beiden Seiten vom Durchschnitt abweichen darf. "
"Bei 4 reicht sie von einem Viertel bis zum Vierfachen."),
FR("De combien le débit peut s'écarter de la moyenne, dans les deux sens. "
"4 le laisse aller du quart au quadruple."),
ES("Cuánto puede alejarse la tasa del promedio, en ambos sentidos. Con 4 "
"va de la cuarta parte al cuádruple."),
PT("Quanto a taxa pode afastar-se da média, nos dois sentidos. Com 4 vai "
"de um quarto ao quádruplo."),
IT("Di quanto la frequenza può scostarsi dalla media, in entrambi i sensi. "
"Con 4 va da un quarto al quadruplo."),
NL("Hoever het tempo van het gemiddelde mag afwijken, beide kanten op. "
"Bij 4 loopt het van een kwart tot vier keer."),
RU("Насколько частота может отходить от средней в обе стороны. При 4 она "
"идёт от четверти до четырёхкратной."),
TR("Hızın ortalamadan iki yöne de ne kadar sapabileceği. 4 olunca dörtte "
"birinden dört katına kadar gider."));
SS_MSG(pano360_output,
EN("Unwrap into (360 video)"),
+15
View File
@@ -551,6 +551,21 @@ SS_MSG(using_bundled_masks,
TR("Fotoğraflarla birlikte gelen maskeler kullanılıyor: {0}"));
// {0} is a whole number of degrees.
SS_MSG(motion_plan,
EN("Motion analysis: {0} frames planned, between {1} and {2} per second."),
JA("動き解析: {0} フレームを予定しました(毎秒 {1}〜{2} フレーム)。"),
ZH_HANS("运动分析:计划取 {0} 帧,每秒 {1} 到 {2} 帧。"),
ZH_HANT("運動分析:計畫取 {0} 影格,每秒 {1} 到 {2} 影格。"),
KO("움직임 분석: {0}개 프레임을 계획했습니다(초당 {1}~{2}장)."),
DE("Bewegungsanalyse: {0} Einzelbilder geplant, zwischen {1} und {2} pro Sekunde."),
FR("Analyse du mouvement : {0} images prévues, entre {1} et {2} par seconde."),
ES("Análisis de movimiento: {0} fotogramas previstos, entre {1} y {2} por segundo."),
PT("Análise de movimento: {0} quadros previstos, entre {1} e {2} por segundo."),
IT("Analisi del movimento: {0} fotogrammi previsti, tra {1} e {2} al secondo."),
NL("Bewegingsanalyse: {0} beelden gepland, tussen {1} en {2} per seconde."),
RU("Анализ движения: запланировано {0} кадров, от {1} до {2} в секунду."),
TR("Hareket incelemesi: {0} kare planlandı, saniyede {1} ile {2} arasında."));
SS_MSG(video_autorotate,
EN("The capture asks to be turned {0} degrees; the frames are written already turned."),
JA("この撮影は {0} 度回転して表示するよう指定されています。フレームは回転済みで書き出されます。"),
+63
View File
@@ -945,6 +945,69 @@ SS_MSG(xh_sync,
TR("çok izli bir dosyayı adım adım birlikte çözerek her izin aynı anları tutmasını sağla "
"(tüm izler için tek keskinlik penceresi); aynı adlı kareler böylece bir rig olur"));
SS_MSG(xh_adaptive,
EN("space the kept frames by how much the view changes rather than by time: "
"more where the camera moves fast or passes close to something, fewer "
"where it turns on the spot. --skip then sets the average"),
JA("残すフレームの間隔を時間ではなく見えの変化量で決めます。速く動いたときや近くの物の"
"そばを通ったときは多く、その場で向きを変えただけのときは少なくなります。--skip は"
"平均値の指定になります"),
ZH_HANS("按画面变化量而不是按时间来安排保留的帧:相机移动快或贴近物体时多取,原地转动时"
"少取。--skip 此时表示平均值"),
ZH_HANT("依畫面變化量而非時間安排保留的影格:相機移動快或貼近物體時多取,原地轉動時少取。"
"--skip 此時表示平均值"),
KO("남길 프레임 간격을 시간이 아니라 시야가 바뀐 정도로 정합니다. 빠르게 움직이거나 "
"가까운 물체를 지날 때는 많이, 제자리에서 돌기만 할 때는 적게 남깁니다. --skip은 "
"평균값이 됩니다"),
DE("die behaltenen Bilder nach der Änderung des Blicks statt nach der Zeit verteilen: "
"mehr, wo die Kamera schnell fährt oder dicht an etwas vorbeikommt, weniger, wo sie "
"sich nur dreht. --skip gibt dann den Durchschnitt an"),
FR("espacer les images conservées selon le changement de vue plutôt que selon le temps : "
"davantage là où la caméra va vite ou frôle un objet, moins là où elle pivote sur "
"place. --skip donne alors la moyenne"),
ES("espaciar los fotogramas conservados según cuánto cambia la vista y no según el tiempo: "
"más donde la cámara va rápido o pasa cerca de algo, menos donde solo gira sobre sí "
"misma. --skip pasa a indicar el promedio"),
PT("espaçar os quadros guardados pela mudança da vista em vez do tempo: mais onde a câmara "
"anda depressa ou passa perto de algo, menos onde apenas gira no lugar. --skip passa a "
"indicar a média"),
IT("distanziare i fotogrammi tenuti in base a quanto cambia la vista anziché al tempo: di "
"più dove la camera va veloce o sfiora qualcosa, di meno dove ruota sul posto. --skip "
"indica allora la media"),
NL("de bewaarde beelden verdelen naar hoeveel het beeld verandert in plaats van naar tijd: "
"meer waar de camera snel gaat of vlak langs iets komt, minder waar hij alleen draait. "
"--skip geeft dan het gemiddelde"),
RU("располагать сохраняемые кадры по изменению вида, а не по времени: чаще там, где камера "
"идёт быстро или проходит близко к предмету, реже там, где она лишь поворачивается на "
"месте. --skip тогда задаёт среднее"),
TR("saklanan kareleri zamana göre değil görüntünün ne kadar değiştiğine göre yerleştir: "
"kamera hızlı giderken ya da bir şeyin yakınından geçerken daha sık, yerinde dönerken "
"daha seyrek. --skip böylece ortalamayı verir"));
SS_MSG(xh_adaptive_range,
EN("how far the adaptive rate may stray from the average, either way "
"(default 4: a quarter of it to four times it)"),
JA("可変レートが平均からどれだけ離れてよいかです(既定 4: 平均の 1/4 から 4 倍まで)"),
ZH_HANS("自适应帧率相对平均值的上下浮动倍数(默认 4:平均值的 1/4 到 4 倍)"),
ZH_HANT("自適應影格率相對平均值的上下浮動倍數(預設 4:平均值的 1/4 到 4 倍)"),
KO("가변 프레임 속도가 평균에서 벗어날 수 있는 배수입니다(기본 4: 평균의 1/4에서 4배)"),
DE("wie weit die angepasste Rate nach beiden Seiten vom Durchschnitt abweichen darf "
"(Vorgabe 4: ein Viertel bis das Vierfache)"),
FR("de combien le débit adaptatif peut s'écarter de la moyenne, dans les deux sens "
"(4 par défaut : du quart au quadruple)"),
ES("cuánto puede alejarse la tasa adaptativa del promedio, en ambos sentidos "
"(4 por defecto: de la cuarta parte al cuádruple)"),
PT("quanto a taxa adaptativa pode afastar-se da média, nos dois sentidos "
"(4 por omissão: de um quarto ao quádruplo)"),
IT("di quanto la frequenza adattiva può scostarsi dalla media, in entrambi i sensi "
"(4 di default: da un quarto al quadruplo)"),
NL("hoever het aangepaste tempo van het gemiddelde mag afwijken, beide kanten op "
"(standaard 4: een kwart tot vier keer)"),
RU("насколько адаптивная частота может отходить от средней в обе стороны "
"(по умолчанию 4: от четверти до четырёхкратной)"),
TR("uyarlanan hızın ortalamadan iki yöne de ne kadar sapabileceği "
"(varsayılan 4: dörtte birinden dört katına)"));
SS_MSG(xh_track,
EN("video track to read; default is every track, written to <out>/cam0, "
"<out>/cam1, ..."),
+48 -4
View File
@@ -205,7 +205,7 @@ struct YuvParams {
};
struct ThumbParams {
vk::DevicePtr out, luma;
uint32_t src_w, src_h, luma_stride, size, flags, shift, groups_per_row;
uint32_t src_w, src_h, luma_stride, out_w, out_h, flags, shift, groups_per_row;
};
struct VarianceParams {
vk::DevicePtr out, thumb;
@@ -289,7 +289,8 @@ struct VideoPipeline::Impl {
// Conversion scratch.
vk::DevicePtr luma_buf = 0, chroma_buf = 0, chroma2_buf = 0;
vk::DevicePtr thumb_buf = 0, metric_buf = 0, rgb_buf = 0;
vk::DevicePtr thumb_buf = 0, metric_buf = 0, rgb_buf = 0, gray_buf = 0;
VkDeviceSize gray_capacity = 0;
VkDeviceSize rgb_capacity = 0;
std::vector<float> metrics;
std::vector<int> metric_pending;
@@ -378,7 +379,8 @@ VideoPipeline::Impl::~Impl() {
for (VkDeviceMemory m : session_mem) vkFreeMemory(dev, m, nullptr);
if (vtimeline) vkDestroySemaphore(dev, vtimeline, nullptr);
if (vpool) vkDestroyCommandPool(dev, vpool, nullptr);
for (vk::DevicePtr p : {luma_buf, chroma_buf, chroma2_buf, thumb_buf, metric_buf, rgb_buf})
for (vk::DevicePtr p : {luma_buf, chroma_buf, chroma2_buf, thumb_buf, metric_buf,
rgb_buf, gray_buf})
if (p) vk::device_free(p);
}
@@ -1310,7 +1312,8 @@ void VideoPipeline::queueSharpness(const FrameHandle& h) {
tp.src_w = (uint32_t)s.fmt.width;
tp.src_h = (uint32_t)s.fmt.height;
tp.luma_stride = s.coded.width;
tp.size = kThumbSize;
tp.out_w = kThumbSize;
tp.out_h = kThumbSize;
tp.flags = s.planeFlags();
tp.shift = (uint32_t)s.planes.shift;
vk::Stream::get().dispatchFlat("video.luma_thumbnail", {}, (int64_t)kThumbSize * kThumbSize,
@@ -1359,6 +1362,47 @@ bool VideoPipeline::Impl::ensureRgb(VkDeviceSize bytes) {
return rgb_buf != 0;
}
bool VideoPipeline::toGray(const FrameHandle& h, int w, int gh,
std::vector<uint8_t>& out, std::string& error) {
Impl& s = *impl_;
if (h.slot < 0 || w <= 0 || gh <= 0) {
error = "toGray: invalid request";
return false;
}
const VkDeviceSize bytes = (VkDeviceSize)w * gh * 4;
if (s.gray_capacity < bytes) {
if (s.gray_buf) vk::device_free(s.gray_buf);
s.gray_buf = vk::device_alloc(bytes, "video.gray");
s.gray_capacity = s.gray_buf ? bytes : 0;
}
if (!s.gray_buf) {
error = "out of device memory for the grey frame buffer";
return false;
}
s.copyPlanes(h.slot);
ThumbParams tp{};
tp.out = s.gray_buf;
tp.luma = s.luma_buf;
tp.src_w = (uint32_t)s.fmt.width;
tp.src_h = (uint32_t)s.fmt.height;
tp.luma_stride = s.coded.width;
tp.out_w = (uint32_t)w;
tp.out_h = (uint32_t)gh;
tp.flags = s.planeFlags();
tp.shift = (uint32_t)s.planes.shift;
vk::Stream::get().dispatchFlat("video.luma_thumbnail", {}, (int64_t)w * gh, 256,
&tp, sizeof(tp), &tp.groups_per_row);
std::vector<float> f((size_t)w * gh);
vk::Stream::get().download(f.data(), s.gray_buf, bytes);
out.resize(f.size());
for (size_t i = 0; i < f.size(); i++)
out[i] = (uint8_t)std::min(255.0f, std::max(0.0f, f[i] + 0.5f));
s.pool[(size_t)h.slot].read_value = 0;
return true;
}
bool VideoPipeline::toImage(const FrameHandle& h, const ConvertOpts& opts, nn::Image& out,
std::string& error) {
Impl& s = *impl_;
+6
View File
@@ -70,6 +70,12 @@ public:
void flushSharpness();
float sharpness(const FrameHandle& h) const;
// A box-filtered grey copy of the frame at `w` x `h`, on the host. What
// motion analysis reads: the GPU does the reduction, so a frame costs a
// dispatch and a few tens of kilobytes over the bus rather than a download.
bool toGray(const FrameHandle& h, int w, int h_out, std::vector<uint8_t>& out,
std::string& error);
// Converts to host RGB. Blocks until the frame is ready.
bool toImage(const FrameHandle& h, const ConvertOpts& opts, nn::Image& out,
std::string& error);
+10 -9
View File
@@ -151,12 +151,13 @@ void yuv_to_rgb(uint3 gid : SV_GroupID, uint3 tid : SV_GroupThreadID,
// written out.
struct ThumbParams {
float* out; // S*S
float* out; // out_w * out_h
uint* luma;
uint src_w;
uint src_h;
uint luma_stride;
uint size; // S
uint out_w;
uint out_h;
uint flags;
uint shift;
uint groups_per_row;
@@ -167,14 +168,14 @@ struct ThumbParams {
void luma_thumbnail(uint3 gid : SV_GroupID, uint3 tid : SV_GroupThreadID,
uniform ThumbParams p) {
uint i = fold_index(gid, tid, p.groups_per_row, uint(WG));
if (i >= p.size * p.size) return;
uint ox = i % p.size;
uint oy = i / p.size;
if (i >= p.out_w * p.out_h) return;
uint ox = i % p.out_w;
uint oy = i / p.out_w;
uint x0 = (ox * p.src_w) / p.size;
uint x1 = max(x0 + 1u, ((ox + 1u) * p.src_w) / p.size);
uint y0 = (oy * p.src_h) / p.size;
uint y1 = max(y0 + 1u, ((oy + 1u) * p.src_h) / p.size);
uint x0 = (ox * p.src_w) / p.out_w;
uint x1 = max(x0 + 1u, ((ox + 1u) * p.src_w) / p.out_w);
uint y0 = (oy * p.src_h) / p.out_h;
uint y1 = max(y0 + 1u, ((oy + 1u) * p.src_h) / p.out_h);
x1 = min(x1, p.src_w);
y1 = min(y1, p.src_h);