It's specific to the harness. Using dynamic context pruning, the budget cuts it off after a select amount of tokens and the budget message tells the model to use subgents to finish whatever it's thinking about
llama.cpp uses a hard cutoff. The agent then does "something" that is specific to the agent's implementation and configuration. It might summarize and then "finish the thought" with a different model, and then resubmit the prompt to the llama.cpp API endpoint with <think>..</think> prefilled. The primary model then infers the remainder of the reply.
llama.cpp does a hard cut off on budget; it can set a reasoning-message as default but the client _can_ set a per message reasoning-message, so it's possible a smart harness could inspect the cut of thoughts and trim and do whatever.
What do you mean by "redirect it to useful output"? Could you give an example? This sounds interesting.