Photograph a solar panel today. Photograph it again in three months. The difference between those two frames is where every defect lives — and it is the one thing no inspection system keeps. This one does.
jgmzrkpa344jwixw7nbulgh2ju0mojcb.lambda-url.us-east-1.on.aws
— the agent described here, running. Upload a photograph of a panel and watch it decide:
200 in 0.58 s, the median of seven warm requests measured
on 15 September 2026. /queue is the approval queue, /assets/{id}
one asset's history, and /traces/{run_id} the full record of any run.
Two of the APIs this runs on do not exist in OpenCV 4:
cv2.ALIKED finds the keypoints in every capture and its stored baseline, and
cv2.LightGlueMatcher matches them — both in the Features module
that replaced Features2D in 5.x. Version 4 ships no learned detector and no
learned matcher, only descriptor distance. Section 06 shows what the difference buys:
0.997 agreement on a pair where classical ORB reaches 0.409.
Solar panels fail slowly. A cell cracks, a hot spot spreads, dust builds up over a season. Every one of those is obvious when you compare against last time, and nearly invisible in a single photograph.
Field inspection today produces photographs, not comparisons. A technician walks the site, shoots what looks wrong, and files a report. The next visit starts from zero — a different technician, a different angle, a different hour of daylight. Nobody compares against last time, because nobody kept last time in a form you can compare against.
So degradation is caught when it becomes visible to a person standing in front of the panel, which is long after it became measurable.
The agent keeps a baseline for every asset it has ever seen. When a new
capture of panel A-114 arrives, it does not analyse that photograph in
isolation — it aligns it against the stored image of that same panel and reports what moved.
That is the whole idea, and it is deliberately narrow. Every award-winning entry in previous editions analyses one frame or one session. The same asset, seen again three months later is empty ground.
The agent is not a classifier with a report attached. At four points in the loop, a number computed by OpenCV changes what happens next — the capture gets rejected, a different detector gets tried, the agent zooms into a region it chose itself, or a human is asked before anything is filed.
Each tool answers one question an inspector would ask, in the order they would ask it.
Only asked when the operator leaves the panel name empty. Every distinctive point in the new photo votes for the stored panel whose points it resembles most. A clear majority names the panel; a split or thin vote ends as unidentified, so the agent asks instead of guessing.
Before analysing anything, check the photograph itself. Measures three things: how sharp the edges are, how many pixels are blown out to pure white or crushed to pure black, and how much of the frame the panel actually fills. A photograph shot from too far away is not a finding — it is a photograph to take again.
A new photo is never shot from the exact angle of the old one. The tool finds distinctive points in both images — corners, cell junctions — and matches them. Hundreds of matches means it is the same panel, and the matches say precisely how to rotate and stretch the new frame onto the stored one. A handful of matches means it is not.
With both frames superimposed, subtract one from the other. Wherever the result is not zero, something moved. Those pixels are grouped into regions, and each region reports where it sits, how large it is, and how strong the change is.
When a region is too small to judge, the agent crops it and measures again at higher resolution. This is the active-perception step: nobody told it where to zoom. The region comes back in full-frame coordinates so the next tool can pick it up.
Each solar defect has a distinct physical signature, and simple measurements separate them. No trained model is involved yet — the thresholds below were read off the fixtures.
| What you see | Brightness | Colour | Verdict |
|---|---|---|---|
| The cell went dark | −28.7 | −6.0 | Cracked cell |
| Bright, but washed out towards white | +29.6 | −19.0 | Hot spot |
| Bright, and shifted to another colour | +54.1 | +46.4 | Delamination |
| Soft change covering most of the frame | +18.4 | −12.8 | Soiling |
None of those five functions returns a verdict. Not one of them says "this photo is bad" or "call a technician". They return numbers only.
That sounds backwards, and it is the most deliberate choice in the codebase. The award's bar
is explicit: the visual evidence must be shown to change what the system does next. If the
threshold were buried inside assess_quality(), there would be nothing to show —
just a function that returned False for reasons of its own.
Keeping the numbers and the thresholds apart means every decision reduces to one line that a judge can read without trusting anybody:
{"input_metric": "inlier_ratio",
"value": 0.038,
"threshold": 0.3,
"branch": "unrecognized_asset"}
"I could not confirm this is panel A-114. Only 4% of the matched points agreed on a single alignment, and this detector needs 30% — the neural one had already failed its own bar of 90%. I am not filing a defect against an asset I cannot identify."
Alignment is the load-bearing step — everything downstream assumes the two frames are the same panel. Here is what it actually scores, run on this hardware, against synthetic panels where the ground truth is known.
| Pair | ORB (classical) | ALIKED + LightGlue |
|---|---|---|
| Same panel, rotated 6° and scaled | 0.409 | 0.997 |
| A different panel | 0.038 | 0.407 |
| Pure noise | 0.020 | 0.000 |
On a correct match the alignment lands within 0.38 pixels of where it should. The neural detector returns 1000 keypoints with 128-dimension descriptors and costs about 1.5 seconds per pair on CPU — OpenCV 5's DNN engine has no GPU support, so the whole system is designed for Graviton from day one.
Three things in that table were not what the code assumed. Each one was found by running it, and each one changed the design.
On the identical pair, ORB scores 0.409 and ALIKED scores 0.997 — and both are correct. A repetitive grid of identical cells makes ORB match cell 3 against cell 7, so its agreement ratio stays low even when its final alignment is accurate to 0.67 pixels. A single shared threshold would make the fallback detector report "asset not recognized" on every panel it ever sees.
crop_and_rescan normalised the contrast of each crop independently. That
pulls the defective crop and the clean crop towards each other and cancels the very
difference being examined — it measured 53 pixels of change where the full frame saw 202.
Normalisation now happens on the whole frame, and the crop is taken afterwards.
A single confidence score blended region size and intensity using two invented constants, then clamped everything above 1% of the frame to the same value — destroying the ordering between a large region and an enormous one, in the exact number meant to trigger the zoom. It was removed; the two raw measurements travel instead.
Aligning a rotated capture leaves black wherever the homography did not reach, and the diff read that border as the largest changed region — 9% of the frame, burying a real cracked cell six times its size under it. The fix warps a validity mask through the same homography, eroded wider than the diff's own blur, so the comparison only ever runs where both frames actually have pixels.
The plan said "blur below 150 means recapture" — then a well-focused capture, warped into the baseline's frame, measured 173, because interpolation eats half the variance of a sharp image (367 before, 173 after). The faint spot painted 22 grey levels deep measured 32.7 after contrast normalisation. Every threshold in the policy was pinned from a run like those, not from the plan's guess.
Everything the agent knows about a panel lives in one place: a single database partition per asset, so its complete history — every inspection, every baseline it ever had — comes back in one query. Images live in object storage, filed by asset and inspection, and every derived artifact (the aligned frame, the validity mask) is stored next to the capture that produced it. A run can be reconstructed from storage alone.
A baseline is never overwritten. When a new capture earns baseline status, the old one is marked superseded and stays — the chain of superseded baselines is the longitudinal record, the thing the project is named after. And memory stores no expiry dates and no verdicts: deciding that a baseline is too old is a threshold, and thresholds belong to the policy, same as everywhere else.
The loop is now running. A language model drives the six instruments over MCP — it chooses the words for the operator and passes the arguments along. What it does not do is decide. After every tool result, plain code compares the numbers against the policy's thresholds, writes the verdict into the run's decision log, and hands the model the verdict it must follow. A model that tries to conclude something the numbers do not support gets its submission rejected, with the mandated branch named in the refusal.
This is the field manual's earlier promise kept: perception reports, policy decides, and the model narrates. Here is a real run — a faint spot, painted 22 grey levels deep, on a capture taken from a different angle:
{"input_metric": "blur_variance", "value": 173.7, "threshold": 100, "branch": "quality_ok"}
{"input_metric": "inlier_ratio", "value": 0.9987, "threshold": 0.9, "branch": "aligned"}
{"input_metric": "mean_delta", "value": 32.68, "threshold": 35, "branch": "crop_and_rescan"}
{"input_metric": "area_ratio", "value": 0.7237, "threshold": 0.02, "branch": "change_confirmed"}
{"input_metric": "score", "value": 0.2431, "threshold": 0.4, "branch": "auto_write"}
The change measured 32.68 against a confirmation bar of 35 — not enough to trust, too much to ignore. So the agent zoomed into the region it chose itself, re-measured with four times the pixels, and the zoomed reading (0.72 of the crop, against a bar of 0.02) confirmed it. A number caused the zoom, and the log proves it.
The fourth branch is the one that stops the machine. When severity crosses its bar, nothing is written: the run parks itself as awaiting approval with everything a reviewer needs, and only an explicit yes commits the inspection and promotes the new baseline. The endpoint is open so a judge can use it without an account, so the gate checks the policy and not an identity: whoever holds the link can resolve it. The approval — or the rejection — is recorded as a decision like any other, with the human as the metric and an anonymous fingerprint of the caller alongside it.
The decision log answers what was decided. The trace answers everything else. Every
run leaves runs/<run_id>/events.json: one event when the run starts, one
span per tool call — the arguments sent, the metrics returned, the milliseconds it took, and
the policy verdict nested inside the span it judged — and one event when the run ends. The
human gate leaves events too: the approval request, and the yes or no that resolved it.
Timestamps come from the server, never the viewer, and the file on disk is the source of
truth — the trace survives a restart and travels as a link.
Here is the run that refused to recognize a foreign panel. The neural matcher voted, the policy sent the agent back for a second opinion, and the classic matcher settled it:
[2026-08-26T22:32:03.493+00:00] run e860a6bbd44a started asset=demo-50ea63 capture=assets/demo-50ea63/capture/capture.png
[2026-08-26T22:32:03.896+00:00] assess_quality (79.7 ms) blur_variance 322.547 >= 100.0 -> quality_ok
[2026-08-26T22:32:05.108+00:00] align_to_baseline (1212.2 ms) inlier_ratio 0.4074 < 0.9 -> retry_classic
[2026-08-26T22:32:05.133+00:00] align_to_baseline (24.2 ms) inlier_ratio 0.0385 < 0.3 -> unrecognized_asset
[2026-08-26T22:32:05.133+00:00] run finished: completed (unrecognized_asset)
The neural matcher scored 0.4074 against its bar of 0.90, so the policy ordered a retry with the classic detector. That retry scored 0.0385 against its own bar of 0.30, and the asset was declared unrecognized. Each next action was caused by the number before it — and the trace proves the causality in a field, not in prose.
The same trace is served over HTTP. GET /traces/<run_id> returns the raw
events as JSON; a browser sending Accept: text/html gets the same run rendered
as a page, with the causal line highlighted on every span. Anything that is not a
well-formed run id — twelve hex characters — is a 404 before the filesystem is consulted,
and the endpoint reads from disk on every request, so the same run always returns the same
bytes. Without a server, python -m services.observability.render <run_id>
prints the text rendering above; the runs root is configurable with
AFTERIMAGE_RUNS_DIR.
Each event carries the sha256 of the event before it, and its own hash over the canonical
form of its fields. GET /traces/<run_id> reports whether the chain
closes, and the page prints it in the footer: edit a metric, a threshold or a branch after
the run, and the verdict names the event it happened in. It is a chain, not a signature —
it catches an edit, a removal or a reordering; it does not catch a trace truncated at the
end, nor anyone who rewrites every hash from the first link.
Everything above happens inside a single container image on one AWS Lambda function. There is no separate inference service, no queue, no orchestrator: the agent loop is a function call made while your upload is still open, which is why the trace is ready by the time the page redirects. The same diagram opens the technical report.
That function sleeps when nobody is using it, and waking it costs 2.34 s — the median of 211 cold starts across the fortnight of logs the group holds. Awake, the function itself answers in 4 ms; the 0.58 s quoted at the top of this page is that plus the round trip to us-east-1, which is the larger share. The EventBridge rule in the diagram exists to keep the awake number the one you get, and the evidence that it works is not the headline ratio — 4.3% of invocations paid an init, but most of those invocations are the warmer's own — it is that no request which paid one was ever an inspection. The longest was 295 ms.
An inspection is the one number no warmer can move: 20.4 s at the median, 27.1 s at worst. About 8 s of it is ALIKED and LightGlue on a CPU and the rest is Gemini deciding, and it is why the page streams the trace instead of spinning. It bills 20.38 seconds at 2 GB, which at the published arm64 rate is $0.0005 a run — in practice nothing, because the whole function burns 2,381 GB-s a fortnight against the 400,000 a month AWS never charges for. The account's bill is image storage, near $0.10 a month. Where each of these figures comes from.
flowchart TB
Browser["Operator browser
no login"]
subgraph deploy["Deploy"]
GHA["GitHub Actions
ubuntu-24.04-arm"]
ECRR["ECR repository
lifecycle keeps the last 5 images"]
CFN["CloudFormation and SAM
infra/template.yaml"]
GHA -->|"OIDC, sub pinned to the
immutable owner and repo IDs"| CFN
GHA -->|"docker buildx, linux/arm64"| ECRR
end
subgraph aws["AWS us-east-1"]
URL["Lambda Function URL
AuthType NONE"]
subgraph fn["One container"]
LWA["Lambda Web Adapter
arm64 Graviton, 2048 MB, 900 s"]
API["FastAPI, serves the whole site"]
LOOP["Agent loop
policy evaluated in code"]
MCPS["MCP server, six perception tools"]
CV["OpenCV 5.0.0 headless
ALIKED and LightGlue ONNX"]
LWA --> API --> LOOP --> MCPS --> CV
end
DDB[("DynamoDB, one table
the asset, its inspections
and its chain of baselines")]
S3B[("S3, captures and traces
objects expire at 180 days")]
LOGS["CloudWatch Logs
kept 30 days"]
EB["EventBridge, every 5 minutes"]
end
Browser -->|"uploads a capture"| URL
URL --> LWA
LOOP <--> DDB
LOOP <--> S3B
fn --> LOGS
EB -->|"a synthetic /health event"| fn
CFN --> fn
ECRR -->|"the image, tagged with the git SHA"| fn
A handful of commands, no cloud account, no credentials. The neural matchers need two ONNX files that are not bundled in the OpenCV wheel; the first command fetches them and checks their published SHA-1.
make weights # 52 MB of ALIKED + LightGlue models
make test # the suite, inside the arm64 container
make demo # the four branches, fired one by one
make dev # the stack up, trace endpoint on :8000
Expect the complete suite to pass on aarch64 with OpenCV 5.0.0. The exact count
changes as coverage grows, so make test is the source of truth. Unit tests use
synthetic solar panels with controlled defects, while evaluation scenarios also include
committed photographs and remain reproducible offline.
One thing worth checking, because it is the honest way to read a green suite: the tests that
exercise the neural path skip themselves when the model files are missing. If
make weights did not run, the suite still passes — with the interesting half
untested.
make demo is the loop's closing argument: four prepared scenarios — a blurred
capture, a foreign panel, a faint spot, a deep crack — each driven through the real MCP
server and the real memory, each ending on a different branch, each printing its full trace
— spans, durations, verdicts, and the human resolution where the gate fires. By default a
scripted driver follows the verdicts so the run is deterministic; with a
Gemini API key, --live puts the model in the driver's seat over the same loop.
The verdicts are computed in code either way.
The loop closes, explains itself, runs in public, has a score, and has the competition report that explains it end to end: the technical report, which is where this manual sends you when you want the numbers rather than the plain-language version, and the three-minute walkthrough if you want neither. The implementation is documented separately for maintainers.
Start with the functional guide, then use the technology stack, architecture, backend guide, and frontend guide for the current system rather than the historical build plans.
Three shorter documents answer the questions this manual deliberately does not: responsible use and limits — what the agent is for, what it is not, and what its numbers do and do not let anyone claim; the security model — a public endpoint by design, what is checked before anything is stored, and what the deploy role can reach; and the use of AI — what the model decides inside the product, what it cannot decide, and what was verified by hand while building it.
| Component | State | Week |
|---|---|---|
| arm64 container, OpenCV 5, CI | built | 1 |
| The five perception tools | built | 2 |
| Memory bank — baselines in S3 and DynamoDB | built | 3 |
| MCP server, agent loop, human gate | built | 4 |
| Per-run event trace, trace endpoint | built | 5 |
| AWS deploy, public endpoint, front end | built | 6 |
| Evaluation set, metrics, failure cases | built | 7 |
| Technical report, both diagrams, video script | built | 8 |
inlier_ratio is not comparable across detectors, so the policy holds one
threshold per detector — 0.90 for the neural pair, 0.30 for ORB — never a shared one.
crop_and_rescan buys measurement precision on a marginal region. It cannot
resolve optical detail the original capture never recorded — no amount of zoom invents
pixels.