A score starts with two records: clean child-facing gameplay with system audio, and a guided review covering access, pricing, offline use, permissions, rewards, and representative learning activities.

The testing process

  1. Sample representative play. We aim for 60–120 seconds per materially different activity and at least three activity types when they exist.
  2. Measure the recording. Deterministic video and audio analysis produces raw metrics; the structured review supplies facts that cannot be read reliably from pixels alone.
  3. Confirm the observations. Draft findings are checked and corrected by a human before scores become canonical.
  4. Calculate versioned scores. Raw evidence, normalized components, formula version, and final scores remain separate so a future calibration can be rerun without re-recording.

The five public scores

Calmness

Gameplay recordings measure motion, visual change, brightness variation, audio intensity and frequency. We also record reward activity and attention prompts. The public score reverses attention intensity: 10 means calmer.

Travel

We test offline function and reliability, preparation burden, dependence on audio, commercial interruptions, and how much content remains accessible.

Content value

We combine content breadth, depth, meaningful variety, and the amount of useful content available without payment.

Learning value

Representative activities are scored for skill demand, agency, depth, and transferability. This measures the interaction—not educational marketing—and does not claim proven learning outcomes.

Commercial experience

We document advertising, purchase pressure in the child flow, visible locked content, pricing transparency, and purchase-model friction.

Exact score formulas

Calmness

10 − Attention Intensity
ComponentWeight
Visual motion25%
Visual change frequency15%
Brightness variability10%
Audio intensity15%
Audio event frequency10%
Reward activity15%
Attention prompts / interruptions10%

Travel

Σ(normalized component × weight)
ComponentWeight
Offline functionality45%
Offline reliability10%
Offline preparation burden15%
Audio independence10%
Child-flow commercial interruptions10%
Accessible content breadth10%

Content value

Σ(normalized component × weight)
ComponentWeight
Content breadth35%
Content depth25%
Meaningful variety20%
Free-access value20%

Learning value

Per activity: (skill demand + agency + depth + transferability) ÷ 8 × 10. App: mean weighted by accessible-experience share.
ComponentWeight
Skill demand25%
Agency25%
Depth25%
Transferability25%

Commercial experience

Σ(normalized component × weight)
ComponentWeight
Third-party advertising25%
Child-flow purchase pressure25%
Locked-content exposure15%
Pricing transparency20%
Purchase-model friction15%

Measurement and normalization

Capture parameters

Video sampling2 frames/second
Analysis width320 px
Changed-pixel luma threshold24
Motion magnitude threshold1.25
Brightness-change threshold18.0
Large visual-change fraction0.28
Silence threshold−45 dB
Audio-event threshold−22 dB
Audio outlier margin8 dB
Audio-event window0.5 seconds
Evidence screenshots3 per category

Raw metric → 0–10

MetricGoodPoorMeaning
Visual motion≤ 0.015≥ 0.18Mean optical-flow motion fraction
Visual changes≤ 2/min≥ 18/minLarge frame-change events
Brightness variability≤ 18≥ 65Luminance standard deviation
Audio intensity≤ −28 LUFS≥ −12 LUFSLess negative is louder
Audio events≤ 4/min≥ 35/minHigh-energy audio events

Values between the good and poor anchors are linearly normalized. These thresholds are provisional for the first ten-app calibration pass.

Observation rubric

How observed answers become component values

ComponentMapping to 0–10
Reward activitynone 0 · light 3 · moderate 6 · heavy 10
Attention promptsno 0 · yes 8
Offline functionalityyes 10 · partial/unreliable 4 · no 0
Offline reliabilityreliable 10 · sometimes fails 5 · always fails 0
Offline preparation burdenno 10 · partial 4 · yes 2
Audio independenceyes 10 · no 2
Commercial interruptionsno 10 · yes 1
Third-party advertisingno 10 · yes 0
Purchase pressureno 10 · yes 1
Locked-content exposureno 10 · yes 3
Pricing transparencyyes 10 · partially hidden option 5 · no 2
Purchase modelone-time/lifetime 10 · subscription 5 · hidden lifetime/subscription front 4
Free accessgenuinely free 10 · meaningful tier 6 · limited demo 4 · trial 3 · effectively paid 1
Meaningful varietylow 3 · moderate 6 · high 8 · excellent 10

For accessible content breadth, 0 items maps to 0 and 12+ maps to 10. Content breadth maps from 1 to 8 activity types; content depth from 3 to 40 items. Unknown always remains null and makes the affected score incomplete.

What the scores do not say

The framework does not label an app addictive, non-addictive, globally safe, unsafe, developmentally beneficial, or proven to produce learning outcomes. Permissions and privacy observations are factual indicators, not a single safety score.

Missing evidence stays unknown. It is never silently converted into “no” or a zero. An incomplete score remains incomplete until the required evidence exists.

See tested apps

Denny's Maze icon

Denny's Maze

The app I built around these priorities

I built Denny's Maze to be optimized for calmness, reliable offline play, and educational value through focused maze-solving. It is my practical attempt to apply the same principles this methodology measures.

View on the App Store