All chapters Part III · Design
Chapter 14

Behavioral Design and Ethics

Every product shapes behaviour whether you intend it or not. The only question is whether you're doing it deliberately, and whose interest you're optimising for. This chapter covers both the mechanics and the line.


The concept

Behavioral design is applying what's known about motivation to product decisions. It's neither inherently good nor inherently manipulative, a streak that helps someone learn a language and an infinite feed that costs them an evening use overlapping mechanics with opposite outcomes.

The distinguishing question is simple: does this mechanic serve the user's stated goal, or does it serve my metric at the user's expense? Ask it about every mechanic you add, and write down the answer.

The motivation stack

Self-determination theory identifies three intrinsic drivers. Products that satisfy them retain without coercion:

DriverWhat it meansHow products satisfy it
AutonomyI chose thisUser-set goals, real customisation, easy exit
CompetenceI'm getting betterVisible progress, appropriate difficulty, feedback
RelatednessI belong / I'm seenCommunity, shared progress, a voice that knows you

Extrinsic motivators, points, badges, leaderboards, streaks, are cheaper to build and weaker over time. Worse, they can displace intrinsic motivation: reward someone for something they already enjoyed and they may enjoy it less once the reward stops. This is the overjustification effect, and it's the reason gamification bolted onto an already-loved product often backfires.

📐 Use extrinsic mechanics as scaffolding for intrinsic motivation, not as a substitute. Points that measure genuine progress are fine. Points invented to create engagement are noise, and users see through them within a week.

Identity beats points

The strongest durable motivator is identity: "I am the kind of person who does this." Products that let users become someone retain far better than products that let users accumulate something.

Weaker (accumulation)Stronger (identity)
"You earned 750 XP""You're now a Level 3 Analyst"
"12-day streak""You've written every day for two weeks"
"You completed 40 lessons""You can hold a basic conversation"

Same underlying data. Completely different meaning. The second column is a claim about the person; the first is a claim about a number.

The mechanics, and their failure modes

MechanicWorks becauseFails when
StreaksLoss aversionHard reset → guilt-churn
Progress barsGoal-gradient effectProgress can go backwards
Variable rewardsUnpredictability sustains attentionIt becomes the whole product
Social proofWe copy othersIt's fabricated
Commitment devicesConsistency biasCoerced rather than chosen
Loss framingLosses loom larger than gainsIt's manufactured or false
Scarcity / urgencyFear of missing outThe scarcity is fake
DefaultsInertiaDefaults serve you, not the user

Defaults deserve special attention. They're the most powerful and least visible lever you have. Most users never change them. Choosing a default is choosing what most people will do, treat it as an ethical decision, not a convenience one.

The line

There's a spectrum, and the boundary is not vague:

 ✅ LEGITIMATE                    ⚠️ GREY                    ❌ DARK PATTERN
 ─────────────────────────────────────────────────────────────────────────
 Real deadline                    Recurring "limited"        Fake countdown
                                  offer                      that resets

 Genuine testimonials             Cherry-picked reviews      Fabricated reviews
                                                             or invented counts

 Making the good path easy        Making the bad path        Making cancellation
                                  slightly duller            hidden or hostile

 Reminder they asked for          Re-engagement they         Notifications
                                  didn't                     designed as bait

 Clear pricing before signup      Price behind a click       Hidden charges,
                                                             pre-ticked add-ons

 Opt-in sharing                   Opt-out sharing            Sharing with no
                                                             real opt-out

The two tests that resolve most cases:

  1. The disclosure test. If you explained this mechanic plainly to the user, would they still be fine with it? "We show a countdown that resets every time you visit", no.
  2. The regret test. After six months, will this user be glad they engaged the way you designed them to? An education app: yes. A slot machine: no.

This is increasingly not just an ethics question. The FTC has taken action on dark patterns and "negative option" billing; the EU's Digital Services Act restricts manipulative interface design; app stores reject fake urgency and misleading subscription flows (Chapter 36). Cancellation flows in particular are under active regulatory scrutiny in multiple jurisdictions.

Vulnerable contexts

If your product touches mental health, finance, addiction, children, or health, the bar rises sharply:


📐 Best practice

Write down what each mechanic is for, and whose interest it serves.

Prefer identity framing over accumulation framing. It costs a rewrite of a label and retains better.

Make progress derive from real value. Points for actions that genuinely help the user; never points for opening the app.

Never let anything decrease. Progress bars that go backwards, levels that drop, streaks that reset to zero, each is a small punishment for imperfect use, and people avoid punishment by avoiding the product.

Give users the controls. Notification frequency, goal size, reminder timing. Autonomy is a retention mechanic, not a concession.

Make cancellation as easy as signup. Same number of clicks, same prominence. It's increasingly a legal requirement and it always builds trust.

Design celebration to be non-blocking. A reward that interrupts the user is the product saying "stop what you're doing and look at me."

Write down the manipulations you refuse, with reasons. It makes the line concrete when you're later under pressure to convert.


💀 Common mistakes

Gamification bolted onto a product nobody loves. Points don't create value; they measure it.

Hard streak resets. The single most self-defeating mechanic in habit products.

Punishing feedback in a self-improvement context. Red X's, error sounds, "you failed", in a product for people trying to improve, punishment reinforces the shame that brought them.

Fake urgency. Countdown timers that reset, "only 3 left" that's always 3. Beyond store rejection, it's a lie in the first minutes of a relationship you want to last.

Fabricated social proof. Invented review counts and testimonials attributed to people who don't exist. Fraud, and trivially disprovable.

Notification bait. Every irrelevant push permanently reduces the value of the channel. Mute is forever.

Defaults that serve you. Pre-ticked marketing consent, auto-renew buried, sharing on by default.

Hostile cancellation. Under regulatory scrutiny and it produces the angriest reviews you will ever receive.

Confusing engagement with value. Time-in-app is a terrible north star for most products. For some, meditation, education, wellbeing, less time can mean more success.


The professional workflow

 1. NAME THE USER'S STATED GOAL. In their words.

 2. FOR EACH MECHANIC, WRITE:
      what it's for · whose interest it serves ·
      what happens to a user who engages imperfectly

 3. RUN THE TWO TESTS
      disclosure — would they be fine if I explained this plainly?
      regret     — will they be glad in six months?

 4. AUDIT FOR PUNISHMENT
      Does anything decrease, reset, or say "behind"? Remove it.

 5. AUDIT DEFAULTS
      Does each default serve the user or me?

 6. CHECK THE EXIT
      Is cancelling as easy as joining?

 7. VULNERABLE-CONTEXT PASS (if applicable)
      safety content free · no punishing mechanics · crisis resources ·
      careful loss framing

 8. WRITE THE REFUSAL LIST — mechanics you won't use, and why

Tools, websites & costs

NeedResourceCost
Dark pattern referencedeceptive.design (Harry Brignull's taxonomy)Free
RegulatoryFTC on dark patterns, EU Digital Services Act summariesFree
Store rulesApp Store Review Guidelines §3.1, §4Free
Behavioural foundationsHooked (Eyal) + his own Indistractable as the counterweight~$30
Self-determination theoryRyan & Deci, summarised widelyFree
Ethical design frameworksCenter for Humane Technology, Ethical OSFree
Habit scienceAtomic Habits (Clear); BJ Fogg's behaviour model~$30

Cost: about $60 of books. This is a thinking chapter, not a tooling one.


Alternatives & trade-offs

Engagement-maximising vs outcome-maximising. Engagement is easier to measure and easier to grow. Outcome is what users actually want and what produces word-of-mouth. Products that optimise engagement at the expense of outcome win short-term metrics and lose reputation, the trajectory of most social feeds.

Intrinsic vs extrinsic. Extrinsic is fast to build and decays; intrinsic is slow and durable. Use extrinsic scaffolding early, then transfer to intrinsic (progress that means something, identity, community).

Social vs solo mechanics. Social motivators are stronger and bring comparison anxiety, moderation load, and a cold-start problem. For wellbeing-adjacent products, social comparison is frequently harmful.

Gentle vs demanding. Demanding products (hard streaks, strict schedules) produce better results for the minority who stick and much higher churn. Gentle products retain more people at lower intensity. Match this to your audience: users who already feel like they're failing need forgiveness; users seeking a challenge want rigour.


Checklist


📓 Case Study: a product that decided not to punish

Project: SOLIS, a self-improvement app whose stated audience was "weak men who want to grow, confidence, success, fighting depression."

That audience makes the ethics load-bearing. A product for people who already feel they're failing cannot afford mechanics that confirm it.

Identity over accumulation, deliberately. The product plan states it as a rule:

Identity beats XP: lessons end with practices and an oath-like completion, not points. Levels exist (Roman ranks) but as chapters of a story, not a grind.

The rank ladder (Tiro → Miles → … → Imperator) is mechanically an XP system. The framing is a military career, which maps onto the user's actual desire, to become someone. The research that informed it noted competitors churning at 2-4 weeks "because XP bars alone don't create attachment."

Nothing ever decreased. No mechanic removed points, reduced a rank, or reset progress. In a product for people who feel like failures, a decreasing number is a small, regular re-injury.

The quiz that refuses to punish. Applied questions after each lesson had explicit rules, a correct pick gets a green tick; a weaker pick gets no red, no green anywhere, just the user's choice outlined and one line explaining why the other is stronger.

This survived a correction. An early build highlighted the better answer in green even when the user hadn't chosen it:

"there is a green tick coming, remove that"

The principle: never show a green "correct" mark next to an answer the user didn't choose. That's the product saying "you were wrong, here's the right one", a test. Dimming the alternative and explaining the distinction is teaching.

Even the haptics carried the rule. iOS offers a harsh notification/ERROR buzz, which is the obvious choice for a wrong answer. It was forbidden in code with a comment: wrong answers get a light, neutral tap. Physical acknowledgement without physical disapproval.

Rewarding the attempt, not just the outcome. Points were awarded for answering a question (4) with a bonus for the stronger choice (+6). If you only reward correctness, users optimise for correctness, they skip questions they might get wrong, which is exactly the behaviour that stops learning.

Safety content was never gated. An explicit rule: "Never gate: the daily first lesson, the streak itself, safety/crisis content." A wellbeing page with crisis resources was built and kept outside the paywall. In the content itself, a deliberate rule held that rest, self-compassion and reaching out are never framed as the weaker option, and the darkest lesson makes telling someone explicitly the stronger move.

That last one is the sharpest example in the chapter. Without that rule, a well-intentioned question like "You're exhausted and behind, push through, or rest?" with "push through" as the stronger answer ships to someone who downloaded a self-improvement app while depressed. That's not a bug; that's harm. It was prevented by writing the exception into the content file where nobody could miss it.

The refusal list. Studying a competitor's funnel (Chapter 10) produced four documented refusals: fake countdown timers ("a fake countdown is not [acceptable]"), fabricated testimonials and invented review counts, a rating prompt fired before the user had used the product, and tracking-SDK attribution. Each was written down with a reason, which is what makes a line hold under later pressure to convert.

⚠️ Deviation: none of it was validated with real users.

Every judgement above is reasoning, not evidence. Whether the forgiving design retains better than a demanding one for this audience is unknown, no analytics, no users (Chapter 38). It's a well-constructed ethical position and an untested product hypothesis, and those are different things.

A second, smaller gap: the streak was designed forgivingly (user-set goal, today can't break it) but the planned streak freeze, the mechanic that most directly prevents guilt-churn, was specified and never built. The principle was right; one implementation was missing.


Lessons

  1. Every product shapes behaviour. The choice is whether you do it deliberately.
  2. Ask whose interest each mechanic serves, and write the answer down.
  3. Identity beats accumulation. "You're now X" retains better than "you earned 750."
  4. Never let anything decrease. Decreasing numbers punish imperfect use.
  5. Reward the attempt, bonus the quality. Reward only correctness and users stop attempting.
  6. Never mark "correct" on an answer they didn't choose. Decide whether you're testing or teaching.
  7. Two tests resolve most cases: would they be fine if you explained it, and will they be glad in six months?
  8. Defaults are ethical decisions. Most users never change them.
  9. Vulnerable contexts raise the bar. Free safety content, no punishing mechanics, careful loss framing, and write the exceptions where nobody can miss them.
  10. Write your refusal list before you need it. Lines hold under pressure only if they were drawn in advance.

Next: Chapter 15: Choosing Your Tech Stack →

This completes Part III, Design.

Useful? Share this chapter