2022 · Released
JanWehner · December 18, 2022 · PC
Nobody has tracked this yet.
Large language models (LLMs) build up models of the world and of tasks leading them to impressive performance on many benchmarks. But how robust are these models against bad data? Motivated by an example where an actively learning LLM is being fed bad data for a task by malicious actors, we propose a benchmark, This Is Fine (TIF), which measures LLMs robustness against such data poisoning. The benchmark takes multiple popular benchmark tasks in NLP, arithmetics, "salient-translation-error-detection" and "phrase-relatedness" and records how the performance of an LLM degrades as it is being fine...
Be the first to rank it.
Be the first to review This Is Fine(-tuning): A benchmark testing LLMs robustness against bad fine-tuning data.