By post-train I presume you mean a finetune? Unless that's wrong (please correct me if so).
I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?
I haven't looked into model architecture people are working with for this stuff too deeply yet but I presume the core idea is fine-tuning a lightweight reasoning-enabled LLM specifically using search as a metric for training?