
Large language models (LLMs) like GPT-4 can identify a personās age, location, gender and income with up to 85 per cent accuracy simply by analysing their posts on social media.
and at ETH Zurich in Switzerland got nine LLMs to pore through a database of Reddit posts and pick up identifying information in the way users wrote.
Staab and Vero randomly selected 1500 profiles of users who engaged on the platform, thenĀ narrowed these down to 520Ā users for which they could confidently identify attributes like a personās place ofĀ birth, their income bracket, gender and location, either inĀ theirĀ profiles or posts.
Advertisement
When given the posting history of those users, some of the LLMs were able to identify many of these attributes with a high degree of accuracy. GPT-4 achieved the highest overall accuracy with 85 per cent, while LlaMA-2-7b, a comparatively low-powered LLM, was the least accurate model with 51 per cent.
āIt tells us that we give a lot ofĀ ourĀ personal information away on the internet without thinking about it,ā says Staab. āMany people would not assume that you can directly infer their age or their location from how they write, butĀ LLMs are quite capable.ā
Sometimes, personal details were explicitly stated in the posts. For example, some users post their income in forums about financial advice. But the AIs also picked upĀ onĀ subtler cues, likeĀ location-specific slang, and could estimate aĀ salary range from aĀ userās profession and location.
Some characteristics were easier for the AIs to discern than others. GPT-4 was 97.8 per cent accurate at guessing gender, but only 62.5 per cent accurate on income.
āWeāre only just beginning to understand how privacy might be affected by use of LLMs,ā says , at the University of Surrey, UK.
arXiv