William Paul HeraldComputer scientist · Independent builder
Menu

Existing local measurements: what they establish

Measurement companion · September 13, 2026. No new runs were performed for this refresh. These are observations from particular configurations, not a controlled comparison of Colibri versions, CPU versus GPU, cloud efficiency or completed-task energy.

Retained test Observation Important boundary
Completed GLM-5.2 weather workflow, September 11 26:02.24 timed turn; two model rounds; one successful weather tool call. Observation retrieval took 224 ms. Activation/setup excluded. Answer contract was explicitly bounded. Earlier attempts failed or timed out. This is one completed task, not ordinary request latency or a hardware comparison.
GLM-5.2 process-start/reuse diagnostic, September 11 First 458-token request: 476.622 seconds; identical repetition: 6.756 seconds. Both deliberately limited to one output token, finish_reason:length. They were diagnostics, not completed-answer acceptance tests. Initial request had 456 new tokens; repetition reused the whole prefix. OS/storage caches were not cleared.
CPU-only GLM-5.2 candidate, September 12 Short readiness reply: 22.122 seconds. New-prefix probe timed out after 300.809 seconds without generated output. CPU execution, RAM budget and pinning changed together. Earlier GPU diagnostic ran longer than this deadline. The timeout is a censored observation, not proof that one arrangement won.
Reduced pin-budget experiment, September 11 New-prefix diagnostic: 537.782 seconds; identical repetition: 7.112 seconds. Matched baseline comparisons did not finish qualification because resource checks refused continued operation. No isolated speedup can be calculated.

The weather reply's required station, measured units, source and observation-age fields were checked. No claim is made that the data were current at final delivery: age was labeled at retrieval. The final answer was an observation, not a forecast.

Existing instrumentation also showed why metric definitions matter. A model profile could omit initial processing while returning a near-zero window; sampled process read transfers could include cached I/O. Neither proves a zero-duration prefill or physical-disk throughput. Separate wall-clock, phase, cache and device measurements are necessary before assigning a cause.

No trustworthy whole-system electricity comparison, cloud baseline, avoided-cloud-usage measurement, adoption study or data-center-capacity effect was established by these tests. The associated article makes none of those outcome claims.