Training an AI for one skill quietly changes its other answers
Post-training a model for one skill nudges its answers on unrelated questions, and that changes how I think about protecting models from copying.
Post-training a model for one skill nudges its answers on unrelated questions, and that changes how I think about protecting models from copying.