Skip to content

On Markov: The Best Move From Here

Published: at 02:32 AM

I realized my life is one giant Markov chain.

I wrote this on X:

Sometimes I feel like I really love strategy games because they’re such a pure thing. There are constraints, trade-offs, and consequences—but they all exist within the game. Nothing spills over into real life. No side effects, no extra baggage. You just understand the rules, make decisions, and try to find the best move. It’s a closed system. That’s probably what makes it feel so pure.

Then I noticed I live the same way. Limited time, limited resources, so every turn I hunt for the minimum-effort path to what I want. Find the bottleneck. Set priorities. Take the highest-leverage move available right now.

Turn after turn. Each move depends only on where I stand.

Table of contents

Open Table of contents

Memoryless

A Markov chain has exactly one rule: the next state depends only on the current state. Not on the path that got you there.

P(next | now, everything before) = P(next | now)

Sounds cold. It’s weirdly freeing.

Sunk cost is mathematically irrelevant. The years already spent, the money already burned, the argument already lost — none of them are inputs. The best move is a function of where you stand, not of how far you walked to get here.

Emotionally, of course, it’s another story, lol.

The past is in the state

Here’s the catch: what counts as “the state”?

Any process turns Markov if you stuff enough history into the state. So the past didn’t disappear. It got compressed into you: skills, savings, body, habits, the people who still pick up when you call, the stuff you never processed.

The past doesn’t matter. Except it’s all in the state.

It only reaches the future through what it left in you. Which is also the only part of it you can still edit.

A chain has no player

A Markov chain is weather. It happens to you. Nobody picks the next state.

What I described in that tweet has a player. That’s a different object: a Markov Decision Process. Same memoryless state, plus two things:

while alive:
    state = observe()
    action = best_move(state)
    reward = world.step(action)

“Find the bottleneck, take the highest-leverage move” is just solving for the policy. Back in my TTA post I said life is an endless game of TTA, and the whole game is one question: where’s the bottleneck — food, rocks, or actions? Same question, now with a formula. The highest-leverage move is the one with the biggest advantage, A(s, a) = Q(s, a) − V(s) — how much better that move is than what you’d usually do from here.

The two things a chain doesn’t have are exactly the two that make it a life: what you can do, and what you want.

Why games feel pure

Back to the tweet. I said strategy games feel pure because they’re closed. Half right. In RL terms, a game hands you a lot for free:

Strategy gameLife
RulesPrinted in the manualUnknown. Learn by playing.
StateOn the boardPartially hidden, even your own
RewardThe win conditionYou write it. It drifts.
EpisodesRestart anytimeOne run. No save file.
Side effectsStay in the boxSpill onto everyone around you

Everyone notices the first two rows. That’s just life being poker, not chess. The row that matters is the third.

The real luxury of a game isn’t clear rules. It’s that someone else wrote the reward function. You never have to ask what winning means. It’s in the rulebook.

Life ships without a win condition. Don’t write one, and the system writes it for you: money, title, followers, output. You optimize it perfectly and still feel the void. In RL terms, that’s a misspecified reward. In Goodhart’s: when a measure becomes a target, it stops being a good measure.

No save file

In a game I take the risky line all the time. Worst case, I reload.

Losing the reload button changes the math more than it sounds. Take this coin flip: heads, your bankroll goes up 50%. Tails, it drops 40%.

Expected value: +5% per flip. Great bet. Now play it a thousand times in a row. A single run doesn’t live on the average; it lives on the geometric mean, √(1.5 × 0.6) ≈ 0.95. About minus five percent, every flip.

The average gets rich. The typical player goes broke.

The fancy name is non-ergodicity. The poker name is bankroll management. EV only counts if you’re still at the table to collect it.

Markov chains have a name for the states you can’t leave: absorbing states. Ruin. A wrecked body. Trust gone for good. Once you’re in, every move leads back to the same place.

So the first rule isn’t “maximize reward.” It’s stay out of absorbing states. Optimize on top of that.

The tunnel

One more knob. Every MDP has a discount factor, γ: how much future reward is worth next to reward right now. γ near 1, you play the long game. γ near 0, only this turn exists.

Scarcity turns γ down. When time and resources are tight, I find the minimum-effort path — and it works. In Scarcity, Mullainathan and Shafir name the side effect: tunneling. The tighter things get, the narrower you see. Important-but-not-urgent falls out of frame. Sleep. Training. Friends. Thinking.

Low γ plus greedy play is the fastest route to a local optimum. Sometimes the best move is the one that looks wrong this turn — because it buys information, or because it repairs the state.

Rest isn’t a pause from the game. It’s a move that changes the state.

From here

So yes, my life is one giant Markov chain. Technically a decision process: one player, no manual, no save file, no win condition unless I write one.

The loop is short. Look at the board. Find the bottleneck. Stay out of states you can’t leave. Write your own reward before the system writes it for you. Take the best move from here.

The past is in the state. The move is mine.