I don't understand why a model has to follow the instructions. Don't get surprised when it shows its true color. Plus, do users (not the researchers) really check whether the response follows the instructions?
Models don't have to follow instructions, but during RLHF that is one of the things they are scored on so a premise of the idea is part of the model, but it's also balanced on accomplishing the end goal. They model may determine, correctly or incorrectly, that your rules suck and do what it thinks is best.
From what I can tell, there are many issues that aren’t off limits to criticize on Chinese social media. In fact, recurring social media complaints are what spurred development of the hotline system.
It’s mainly complaints that are considered sensitive or destabilizing that are suppressed. This should sound familiar to those of us in the West. Germany actually goes farther by directly funding left-wing protest groups, as these are not considered destabilizing.
Creating accounts should be allowed, but using an account could require age check.
People should be able to create an account at birth. Then when they grow up, they are ready to use the account. This way proves that the account owner is at least as old as the account.
Don't fight with AI. People who reported increased throughput don't verify AI output. Programs which are simply AI wrappers don't verify the output. If a serious programmer starts to design an algorithm that verifies AI output, the progress starts to slow down because you are doing something not needed if you don't use AI.
reply