We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is a useful tool for the conversations going on about this.
But that was about transferring from a finetuned model to another finetune of the same base model, good to see more evidence that it doesn't transfer cleanly across different base models in a more realistic scenario than an "owl-loving model"