afterimage OpenCV AI Competition 2026 Agentic Vision path

An inspection agent
that remembers.

Photograph a solar panel today. Photograph it again in three months. The difference between those two frames is where every defect lives — and it is the one thing no inspection system keeps. This one does.

Try it — live, no login

jgmzrkpa344jwixw7nbulgh2ju0mojcb.lambda-url.us-east-1.on.aws — the agent described here, running. Upload a photograph of a panel and watch it decide: 200 in 0.58 s, the median of seven warm requests measured on 15 September 2026. /queue is the approval queue, /assets/{id} one asset's history, and /traces/{run_id} the full record of any run.

Needs OpenCV 5

Two of the APIs this runs on do not exist in OpenCV 4: cv2.ALIKED finds the keypoints in every capture and its stored baseline, and cv2.LightGlueMatcher matches them — both in the Features module that replaced Features2D in 5.x. Version 4 ships no learned detector and no learned matcher, only descriptor distance. Section 06 shows what the difference buys: 0.997 agreement on a pair where classical ORB reaches 0.409.

01

The problem

Solar panels fail slowly. A cell cracks, a hot spot spreads, dust builds up over a season. Every one of those is obvious when you compare against last time, and nearly invisible in a single photograph.

Field inspection today produces photographs, not comparisons. A technician walks the site, shoots what looks wrong, and files a report. The next visit starts from zero — a different technician, a different angle, a different hour of daylight. Nobody compares against last time, because nobody kept last time in a form you can compare against.

So degradation is caught when it becomes visible to a person standing in front of the panel, which is long after it became measurable.

02

What makes this different

The agent keeps a baseline for every asset it has ever seen. When a new capture of panel A-114 arrives, it does not analyse that photograph in isolation — it aligns it against the stored image of that same panel and reports what moved.

That is the whole idea, and it is deliberately narrow. Every award-winning entry in previous editions analyses one frame or one session. The same asset, seen again three months later is empty ground.

03

How it decides

The agent is not a classifier with a report attached. At four points in the loop, a number computed by OpenCV changes what happens next — the capture gets rejected, a different detector gets tried, the agent zooms into a region it chose itself, or a human is asked before anything is filed.

every step emitted as a span capture of a known asset assess_quality() align_to_baseline() diff_against_memory() classify_severity() write_memory() Is this photo usable? Same panel? How does it fit? What changed? How bad, and what kind? new baseline, event, trace blur_variance inlier_ratio area_ratio score Request recapture Retry with another detector Crop and rescan Request human approval “closer”, “less backlight” or flag the asset as unrecognized the agent picks where to look closer before anything is filed re-enters
Each amber branch is a point where a measured value changes the agent's next move. The metric that triggers the branch is named on the arrow — that name is also the key it carries into the trace. The zoom is the only branch that feeds back: the agent decides where to look closer and re-enters the loop with a tighter frame.
04

The six instruments

Each tool answers one question an inspector would ask, in the order they would ask it.

  1. 01

    Which panel is this?

    identify_asset()

    Only asked when the operator leaves the panel name empty. Every distinctive point in the new photo votes for the stored panel whose points it resembles most. A clear majority names the panel; a split or thin vote ends as unidentified, so the agent asks instead of guessing.

  2. 02

    Is this photo usable?

    assess_quality()

    Before analysing anything, check the photograph itself. Measures three things: how sharp the edges are, how many pixels are blown out to pure white or crushed to pure black, and how much of the frame the panel actually fills. A photograph shot from too far away is not a finding — it is a photograph to take again.

  3. 03

    Is this the same panel, and how does it fit?

    align_to_baseline()

    A new photo is never shot from the exact angle of the old one. The tool finds distinctive points in both images — corners, cell junctions — and matches them. Hundreds of matches means it is the same panel, and the matches say precisely how to rotate and stretch the new frame onto the stored one. A handful of matches means it is not.

  4. 04

    What changed?

    diff_against_memory()

    With both frames superimposed, subtract one from the other. Wherever the result is not zero, something moved. Those pixels are grouped into regions, and each region reports where it sits, how large it is, and how strong the change is.

  5. 05

    I want to see that closer.

    crop_and_rescan()

    When a region is too small to judge, the agent crops it and measures again at higher resolution. This is the active-perception step: nobody told it where to zoom. The region comes back in full-frame coordinates so the next tool can pick it up.

  6. 06

    How bad is it, and what kind?

    classify_severity()

    Each solar defect has a distinct physical signature, and simple measurements separate them. No trained model is involved yet — the thresholds below were read off the fixtures.

What you see Brightness Colour Verdict
The cell went dark −28.7 −6.0 Cracked cell
Bright, but washed out towards white +29.6 −19.0 Hot spot
Bright, and shifted to another colour +54.1 +46.4 Delamination
Soft change covering most of the frame +18.4 −12.8 Soiling
05

Perception decides nothing

None of those five functions returns a verdict. Not one of them says "this photo is bad" or "call a technician". They return numbers only.

That sounds backwards, and it is the most deliberate choice in the codebase. The award's bar is explicit: the visual evidence must be shown to change what the system does next. If the threshold were buried inside assess_quality(), there would be nothing to show — just a function that returned False for reasons of its own.

Keeping the numbers and the thresholds apart means every decision reduces to one line that a judge can read without trusting anybody:

{"input_metric": "inlier_ratio",
 "value": 0.038,
 "threshold": 0.3,
 "branch": "unrecognized_asset"}
In plain words

"I could not confirm this is panel A-114. Only 4% of the matched points agreed on a single alignment, and this detector needs 30% — the neural one had already failed its own bar of 90%. I am not filing a defect against an asset I cannot identify."

06

Measured, not promised

Alignment is the load-bearing step — everything downstream assumes the two frames are the same panel. Here is what it actually scores, run on this hardware, against synthetic panels where the ground truth is known.

Pair ORB (classical) ALIKED + LightGlue
Same panel, rotated 6° and scaled 0.409 0.997
A different panel 0.038 0.407
Pure noise 0.020 0.000

On a correct match the alignment lands within 0.38 pixels of where it should. The neural detector returns 1000 keypoints with 128-dimension descriptors and costs about 1.5 seconds per pair on CPU — OpenCV 5's DNN engine has no GPU support, so the whole system is designed for Graviton from day one.

07

What measuring changed

Three things in that table were not what the code assumed. Each one was found by running it, and each one changed the design.

  1. The two detectors cannot share a threshold

    On the identical pair, ORB scores 0.409 and ALIKED scores 0.997 — and both are correct. A repetitive grid of identical cells makes ORB match cell 3 against cell 7, so its agreement ratio stays low even when its final alignment is accurate to 0.67 pixels. A single shared threshold would make the fallback detector report "asset not recognized" on every panel it ever sees.

  2. The zoom was erasing what it zoomed into

    crop_and_rescan normalised the contrast of each crop independently. That pulls the defective crop and the clean crop towards each other and cancels the very difference being examined — it measured 53 pixels of change where the full frame saw 202. Normalisation now happens on the whole frame, and the crop is taken afterwards.

  3. "Confidence" was a decision wearing a measurement's clothes

    A single confidence score blended region size and intensity using two invented constants, then clamped everything above 1% of the frame to the same value — destroying the ordering between a large region and an enormous one, in the exact number meant to trigger the zoom. It was removed; the two raw measurements travel instead.

  4. The warp's black border was the biggest "change" in the frame

    Aligning a rotated capture leaves black wherever the homography did not reach, and the diff read that border as the largest changed region — 9% of the frame, burying a real cracked cell six times its size under it. The fix warps a validity mask through the same homography, eroded wider than the diff's own blur, so the comparison only ever runs where both frames actually have pixels.

  5. The thresholds moved the moment they were measured

    The plan said "blur below 150 means recapture" — then a well-focused capture, warped into the baseline's frame, measured 173, because interpolation eats half the variance of a sharp image (367 before, 173 after). The faint spot painted 22 grey levels deep measured 32.7 after contrast normalisation. Every threshold in the policy was pinned from a run like those, not from the plan's guess.

08

The memory it keeps

Everything the agent knows about a panel lives in one place: a single database partition per asset, so its complete history — every inspection, every baseline it ever had — comes back in one query. Images live in object storage, filed by asset and inspection, and every derived artifact (the aligned frame, the validity mask) is stored next to the capture that produced it. A run can be reconstructed from storage alone.

A baseline is never overwritten. When a new capture earns baseline status, the old one is marked superseded and stays — the chain of superseded baselines is the longitudinal record, the thing the project is named after. And memory stores no expiry dates and no verdicts: deciding that a baseline is too old is a threshold, and thresholds belong to the policy, same as everywhere else.

09

Who decides what

The loop is now running. A language model drives the six instruments over MCP — it chooses the words for the operator and passes the arguments along. What it does not do is decide. After every tool result, plain code compares the numbers against the policy's thresholds, writes the verdict into the run's decision log, and hands the model the verdict it must follow. A model that tries to conclude something the numbers do not support gets its submission rejected, with the mandated branch named in the refusal.

This is the field manual's earlier promise kept: perception reports, policy decides, and the model narrates. Here is a real run — a faint spot, painted 22 grey levels deep, on a capture taken from a different angle:

{"input_metric": "blur_variance",  "value": 173.7,  "threshold": 100,  "branch": "quality_ok"}
{"input_metric": "inlier_ratio",   "value": 0.9987, "threshold": 0.9,  "branch": "aligned"}
{"input_metric": "mean_delta",     "value": 32.68,  "threshold": 35,   "branch": "crop_and_rescan"}
{"input_metric": "area_ratio",     "value": 0.7237, "threshold": 0.02, "branch": "change_confirmed"}
{"input_metric": "score",          "value": 0.2431, "threshold": 0.4,  "branch": "auto_write"}
Read the third line

The change measured 32.68 against a confirmation bar of 35 — not enough to trust, too much to ignore. So the agent zoomed into the region it chose itself, re-measured with four times the pixels, and the zoomed reading (0.72 of the crop, against a bar of 0.02) confirmed it. A number caused the zoom, and the log proves it.

The fourth branch is the one that stops the machine. When severity crosses its bar, nothing is written: the run parks itself as awaiting approval with everything a reviewer needs, and only an explicit yes commits the inspection and promotes the new baseline. The endpoint is open so a judge can use it without an account, so the gate checks the policy and not an identity: whoever holds the link can resolve it. The approval — or the rejection — is recorded as a decision like any other, with the human as the metric and an anonymous fingerprint of the caller alongside it.

10

The trace

The decision log answers what was decided. The trace answers everything else. Every run leaves runs/<run_id>/events.json: one event when the run starts, one span per tool call — the arguments sent, the metrics returned, the milliseconds it took, and the policy verdict nested inside the span it judged — and one event when the run ends. The human gate leaves events too: the approval request, and the yes or no that resolved it. Timestamps come from the server, never the viewer, and the file on disk is the source of truth — the trace survives a restart and travels as a link.

Here is the run that refused to recognize a foreign panel. The neural matcher voted, the policy sent the agent back for a second opinion, and the classic matcher settled it:

[2026-08-26T22:32:03.493+00:00] run e860a6bbd44a started  asset=demo-50ea63  capture=assets/demo-50ea63/capture/capture.png
[2026-08-26T22:32:03.896+00:00] assess_quality (79.7 ms)  blur_variance 322.547 >= 100.0 -> quality_ok
[2026-08-26T22:32:05.108+00:00] align_to_baseline (1212.2 ms)  inlier_ratio 0.4074 < 0.9 -> retry_classic
[2026-08-26T22:32:05.133+00:00] align_to_baseline (24.2 ms)  inlier_ratio 0.0385 < 0.3 -> unrecognized_asset
[2026-08-26T22:32:05.133+00:00] run finished: completed (unrecognized_asset)
Read the two alignment lines

The neural matcher scored 0.4074 against its bar of 0.90, so the policy ordered a retry with the classic detector. That retry scored 0.0385 against its own bar of 0.30, and the asset was declared unrecognized. Each next action was caused by the number before it — and the trace proves the causality in a field, not in prose.

The same trace is served over HTTP. GET /traces/<run_id> returns the raw events as JSON; a browser sending Accept: text/html gets the same run rendered as a page, with the causal line highlighted on every span. Anything that is not a well-formed run id — twelve hex characters — is a 404 before the filesystem is consulted, and the endpoint reads from disk on every request, so the same run always returns the same bytes. Without a server, python -m services.observability.render <run_id> prints the text rendering above; the runs root is configurable with AFTERIMAGE_RUNS_DIR.

The trace is hash-chained

Each event carries the sha256 of the event before it, and its own hash over the canonical form of its fields. GET /traces/<run_id> reports whether the chain closes, and the page prints it in the footer: edit a metric, a threshold or a branch after the run, and the verdict names the event it happened in. It is a chain, not a signature — it catches an edit, a removal or a reordering; it does not catch a trace truncated at the end, nor anyone who rewrites every hash from the first link.

11

The cloud it runs in

Everything above happens inside a single container image on one AWS Lambda function. There is no separate inference service, no queue, no orchestrator: the agent loop is a function call made while your upload is still open, which is why the trace is ready by the time the page redirects. The same diagram opens the technical report.

That function sleeps when nobody is using it, and waking it costs 2.34 s — the median of 211 cold starts across the fortnight of logs the group holds. Awake, the function itself answers in 4 ms; the 0.58 s quoted at the top of this page is that plus the round trip to us-east-1, which is the larger share. The EventBridge rule in the diagram exists to keep the awake number the one you get, and the evidence that it works is not the headline ratio — 4.3% of invocations paid an init, but most of those invocations are the warmer's own — it is that no request which paid one was ever an inspection. The longest was 295 ms.

An inspection is the one number no warmer can move: 20.4 s at the median, 27.1 s at worst. About 8 s of it is ALIKED and LightGlue on a CPU and the rest is Gemini deciding, and it is why the page streams the trace instead of spinning. It bills 20.38 seconds at 2 GB, which at the published arm64 rate is $0.0005 a run — in practice nothing, because the whole function burns 2,381 GB-s a fortnight against the 400,000 a month AWS never charges for. The account's bill is image storage, near $0.10 a month. Where each of these figures comes from.

flowchart TB
    Browser["Operator browser
no login"] subgraph deploy["Deploy"] GHA["GitHub Actions
ubuntu-24.04-arm"] ECRR["ECR repository
lifecycle keeps the last 5 images"] CFN["CloudFormation and SAM
infra/template.yaml"] GHA -->|"OIDC, sub pinned to the
immutable owner and repo IDs"| CFN GHA -->|"docker buildx, linux/arm64"| ECRR end subgraph aws["AWS us-east-1"] URL["Lambda Function URL
AuthType NONE"] subgraph fn["One container"] LWA["Lambda Web Adapter
arm64 Graviton, 2048 MB, 900 s"] API["FastAPI, serves the whole site"] LOOP["Agent loop
policy evaluated in code"] MCPS["MCP server, six perception tools"] CV["OpenCV 5.0.0 headless
ALIKED and LightGlue ONNX"] LWA --> API --> LOOP --> MCPS --> CV end DDB[("DynamoDB, one table
the asset, its inspections
and its chain of baselines")] S3B[("S3, captures and traces
objects expire at 180 days")] LOGS["CloudWatch Logs
kept 30 days"] EB["EventBridge, every 5 minutes"] end Browser -->|"uploads a capture"| URL URL --> LWA LOOP <--> DDB LOOP <--> S3B fn --> LOGS EB -->|"a synthetic /health event"| fn CFN --> fn ECRR -->|"the image, tagged with the git SHA"| fn
The deploy carries no stored credentials: GitHub federates over OIDC into a role whose trust is pinned to numeric owner and repository IDs, and whose grant reaches only this stack's own repository, stack, function, table, bucket, log group and warmer rule. Nothing else in the account is reachable from a build.
12

Run it yourself

A handful of commands, no cloud account, no credentials. The neural matchers need two ONNX files that are not bundled in the OpenCV wheel; the first command fetches them and checks their published SHA-1.

make weights   # 52 MB of ALIKED + LightGlue models
make test      # the suite, inside the arm64 container
make demo      # the four branches, fired one by one
make dev       # the stack up, trace endpoint on :8000

Expect the complete suite to pass on aarch64 with OpenCV 5.0.0. The exact count changes as coverage grows, so make test is the source of truth. Unit tests use synthetic solar panels with controlled defects, while evaluation scenarios also include committed photographs and remain reproducible offline.

One thing worth checking, because it is the honest way to read a green suite: the tests that exercise the neural path skip themselves when the model files are missing. If make weights did not run, the suite still passes — with the interesting half untested.

make demo is the loop's closing argument: four prepared scenarios — a blurred capture, a foreign panel, a faint spot, a deep crack — each driven through the real MCP server and the real memory, each ending on a different branch, each printing its full trace — spans, durations, verdicts, and the human resolution where the gate fires. By default a scripted driver follows the verdicts so the run is deterministic; with a Gemini API key, --live puts the model in the driver's seat over the same loop. The verdicts are computed in code either way.

13

Where the project stands

The loop closes, explains itself, runs in public, has a score, and has the competition report that explains it end to end: the technical report, which is where this manual sends you when you want the numbers rather than the plain-language version, and the three-minute walkthrough if you want neither. The implementation is documented separately for maintainers.

Start with the functional guide, then use the technology stack, architecture, backend guide, and frontend guide for the current system rather than the historical build plans.

Three shorter documents answer the questions this manual deliberately does not: responsible use and limits — what the agent is for, what it is not, and what its numbers do and do not let anyone claim; the security model — a public endpoint by design, what is checked before anything is stored, and what the deploy role can reach; and the use of AI — what the model decides inside the product, what it cannot decide, and what was verified by hand while building it.

ComponentStateWeek
arm64 container, OpenCV 5, CIbuilt1
The five perception toolsbuilt2
Memory bank — baselines in S3 and DynamoDBbuilt3
MCP server, agent loop, human gatebuilt4
Per-run event trace, trace endpointbuilt5
AWS deploy, public endpoint, front endbuilt6
Evaluation set, metrics, failure casesbuilt7
Technical report, both diagrams, video scriptbuilt8
14

Known limits

  • Defect classification is a threshold heuristic over OpenCV features, not a trained classifier. It is measured on 29 scenarios, eighteen of them real photographs of solar modules: macro F1 0.8542, with precision at least 0.75 on all four defect classes. Most of the suite is photographic — but every lesion in that set was injected by us, so this measures robustness on real texture, not field detection rates.
  • inlier_ratio is not comparable across detectors, so the policy holds one threshold per detector — 0.90 for the neural pair, 0.30 for ORB — never a shared one.
  • Frame coverage is estimated from the bounding box of the largest edge contour, and week 7 showed why that check still ships disabled: across healthy real photographs it reads anywhere from 0.0032 to 0.6357, a two-hundred-fold spread on images that are all perfectly usable. No single default separates a good capture from a badly framed one, so it stays a per-site setting rather than a promise.
  • The agent scores 0.8621 on picking the right branch, and it never once asked a human to look at something that did not warrant it. It fails in four ways worth knowing: a half-framed photo gets rejected for the wrong reason — on a generated panel and on a photograph alike — a very bright defect can trip the exposure check before anyone looks at the defect itself, a small but serious defect can be filed automatically instead of escalated, because severity scores how much changed and ignores what its own classifier just called it, and soiling on a photograph is labelled a hot spot, because the rule that recognises dirt asks whether the change covers a quarter of the frame and real dust breaks into fragments that never do.
  • crop_and_rescan buys measurement precision on a marginal region. It cannot resolve optical detail the original capture never recorded — no amount of zoom invents pixels.