

See https://github.com/JJJHolscher/alignment_jam_2 We investigate a recent model editing technique for large language models called Rank-One Model Editing (ROME). ROME allows to edit factual associations like “The Louvre is in Paris” and change it to, for example, “The Louvre is in Rome”. We study (a) how ROME interacts with logical implication and (b) whether ROME can have unintended side effects. Regarding (a), we find that ROME (as expected) does not respect logical implication for symmetric ...

Tap to see all activity →
No reviews for Model editing hazards at the example of ROME yet
Be the first to write one.