Experiments

Questions we're testing -- honestly.

Every experiment below is labeled Proposed -- none has actually been run yet. When one runs for real, its result replaces the placeholder, never the other way around.

Human vs AI ExperimentsPROPOSED

Can a viewer tell which video thumbnail was AI-selected vs. human-selected?

Question
Does Project KAI's automated thumbnail-ranking system pick thumbnails a human audience actually prefers?
Hypothesis
Automated ranking will perform comparably to, but not better than, a human editor's final pick -- since the pipeline's own Thumbnail Agent explicitly routes final selection to human review today (see /agents).
Method
Show paired thumbnail candidates (auto-ranked top pick vs. human-selected pick) to a small panel and record preference, without revealing which is which.
Result
Not yet run.
Limitations
Proposed only. No panel has been recruited and no data has been collected.
Your reactions only -- not shared or stored anywhere else.
AI Writing ExperimentsPROPOSED

Does a shorter video script hook improve retention more than a longer one?

Question
Within Project KAI's real production pipeline, does hook length correlate with viewer retention?
Hypothesis
Shorter hooks (under 10 seconds) will show higher early retention, based on general short-form video convention -- untested against KAI's own real data so far.
Method
Once real YouTube Analytics access is authorized (currently NOT_CONNECTED -- see /analytics), compare retention curves for videos grouped by hook length.
Result
Not yet run -- blocked on Analytics Agent activation.
Limitations
Requires real analytics access that doesn't exist yet; cannot be run today.
Your reactions only -- not shared or stored anywhere else.
AI Agent ExperimentsPROPOSED

How often does a proposed Developer Agent change require human correction?

Question
If the currently-PLANNED Developer Agent (see /agents) proposed code changes for human review, how often would a human need to correct rather than simply approve them?
Hypothesis
Early proposals will need frequent correction, improving over time as the agent's lesson-lookup (developer_memory/) accumulates more real decisions.
Method
Track approve/correct/reject rates once the Developer Agent exists and is run in its designed sandbox-test-benchmark-approve loop.
Result
Not yet run -- the Developer Agent itself does not exist yet.
Limitations
Fully blocked on an unbuilt agent; listed here to make the eventual evaluation plan explicit in advance.
Your reactions only -- not shared or stored anywhere else.
AI Image ExperimentsPROPOSED

Does avoiding duplicate visual backgrounds actually reduce viewer drop-off?

Question
Project KAI's Visual Agent already checks every scene against prior videos for duplicate backgrounds (see /agents) -- does this measurably help retention, or is it invisible to viewers?
Hypothesis
The effect is small per-video but compounds across a channel's catalog as repeat viewers accumulate exposure.
Method
Compare retention on videos with zero duplicate-flagged scenes vs. videos where a duplicate was allowed through, once real analytics access exists.
Result
Not yet run -- blocked on Analytics Agent activation.
Limitations
Requires real analytics access that doesn't exist yet.
Your reactions only -- not shared or stored anywhere else.
Automation ExperimentsPROPOSED

Would automatic ledger writes catch more real lessons than manual ones?

Question
The developer_memory/ ledger is real and actively used, but written manually today (see /agents' Memory Agent entry). Would automating writes as a side effect of other agents' actions capture more, not just faster?
Hypothesis
Automated writes would increase volume but might reduce signal quality without a human filter deciding what's actually worth recording.
Method
Once other agents exist to trigger automatic writes, compare a sample against manually-curated entries for the same time period for redundancy and usefulness.
Result
Not yet run -- depends on other unbuilt agents.
Limitations
Blocked on multiple other PLANNED agents existing first.
Your reactions only -- not shared or stored anywhere else.
Prompt ExperimentsPROPOSED

Does asking KAI's (future) Comment Assistant for 'shorter' vs 'more detailed' actually change reply quality?

Question
Once a Comment Assistant exists (see PROJECT_KAI_MEDIA_NETWORK_ARCHITECTURE.md's design-only architecture), do its stated reply-tone modes produce meaningfully different, still-accurate replies?
Hypothesis
Tone modes will change length and phrasing reliably but should not change factual content -- a real risk worth testing before ever allowing auto-suggested replies near production comments.
Method
Once built, generate replies to a fixed comment set across all six tone modes and have a human reviewer check factual consistency across all six.
Result
Not yet run -- the Comment Assistant does not exist yet.
Limitations
Fully blocked on an unbuilt system; listed to make the safety-testing plan explicit before that system is ever built.
Your reactions only -- not shared or stored anywhere else.

Discussion

In Development

Comments are not connected to a backend yet -- nothing you type below is saved, sent, or visible to anyone else. This is the real, working interface design; persistence is a future, separately-decided backend project.

Which of these experiments would you want run first?

Not saved anywhere yet -- see note above.

No comments yet -- there's nowhere for them to live. This space is ready for a real comment (name, body, reply, like, report) the moment a backend is connected.

No experiment above has been performed -- every "Result" field says so honestly. This page exists to make the evaluation plan explicit before any system it depends on is built.