step 1, identify high value users by net worth, citation count, or number of followers
step 2, select all prompts by high value users
step 3, invest 10 billion thinking tokens in modeling an objective for each user
step 4, build an RL environment for each user
step 5, rollout 10 billion tokens per environment
step 6, train on resulting traces