option
Home
Flash News
Content
RoyLopez
RoyLopez
April 20, 2026

A Nature study reveals large language models can transmit hidden behavioral traits like biases or unsafe preferences through seemingly neutral data such as number sequences or code, a phenomenon termed subliminal learning. This compromises common model distillation techniques, allowing risks to propagate undetected through model supply chains. Current safety evaluations focused on semantic content fail to identify these non‑semantic signals, urging a shift toward auditing model weights rather than just outputs.

A Nature study reveals large language models can transmit hidden behavioral traits like biases or unsafe preferences through seemingly neutral data such as number sequences or code, a phenomenon termed subliminal learning. This compromises common model distillation techniques, allowing risks to propagate undetected through model supply chains. Current safety evaluations focused on semantic content fail to identify these non‑semantic signals, urging a shift toward auditing model weights rather than just outputs.
Comments (0)
0/300
OR