What is DarkBert?

DarkBERT is a new AI model specifically trained using data from the darknet.

DarkBERT is a new AI model specifically trained on data from the darknet. Unlike large language models such as ChatGPT and Google Bard, which were trained on data from the open web, the developers of DarkBERT used exclusively darknet data for training. More precisely, DarkBERT was trained on data from hackers, cybercriminals, and other fraudsters.

DarkBERT is based on the RoBERTa architecture, an AI method developed by Facebook researchers in 2019. RoBERTa is a "robustly optimized method for preprocessing Natural Language Processing (NLP) systems" that improves upon BERT (Bidirectional Encoder Representations from Transformers), released by Google in 2018. A team of researchers from South Korea scanned the Tor network to gather data for training this comprehensive language model. By feeding RoBERTa with data from the darknet over a period of almost 16 days, the researchers were able to develop DarkBERT.

Despite the unusual origin of its training data, DarkBERT has already outperformed other major language models. The researchers currently have no plans to make DarkBERT publicly available, but are accepting requests for academic use. DarkBERT will likely provide law enforcement and researchers with a better understanding of the darknet as a whole.

DarkBERT could be the future of AI models that are trained in a specific area to make them more specialized. Given its current popularity, it wouldn't be surprising to see similar AI models developed in this way in the future.

Es wurde kein Alt-Text für dieses Bild angegeben.
DarkBERT: Illustration of the pretraining process and evaluation scenarios (Image: DarkBERT: A Language Model for the Dark Side of the Internet)

Why is a DarkBERT even necessary?

In the context of cybersecurity and law enforcement, DarkBERT represents a remarkable tool. It has proven its capabilities in darknet tests, slightly outperforming established models like BERT (a now somewhat outdated model) compared to more powerful transformer models like GPT, thanks to its domain knowledge. DarkBERT, a retrained version of RoBERTa, was trained over two weeks using two different datasets: one with raw crawled data and the other with processed data.

DarkBERT's primary target audience, however, is not cybercriminals, but law enforcement agencies and cybersecurity organizations that scour the darknet to combat internet crime. According to available research, the predominant themes on the darknet are fraud and data theft, and the darknet is also used for anonymous discussions within organized crime.

What is the Darknet?

It is important to note that the Darknet or Deep Web is an area of ​​the internet that conventional search engines like Google do not access and is generally inaccessible to the average user, as special software is required.

While DarkBERT can be an effective tool in the fight against cybercrime, the possibilities of anonymous web browsing are also of interest to many other people, especially those who value their privacy and do not want to make their data available to the large technology companies that have made data collection and personalized advertising their business model. Journalists, dissidents, and the politically persecuted, for example, use the darknet to access regionally blocked and censored content.

Why DarkBERT makes sense but is inaccessible

Overall, DarkBERT is a versatile and powerful tool that not only helps combat internet crime but can also contribute to a better understanding of the Darknet and the activities that take place there.

There are some notable advantages to DarkBERT:

  • It has the ability to identify websites that offer ransomware or disclose sensitive data.
  • It can search various forums on the darknet and draw attention to illegal information exchanges.
  • Despite being trained on Darknet data, DarkBERT has already outperformed other major language models.

Despite these positive aspects, there are also concerns regarding the use of DarkBERT:

  • Because DarkBERT was trained on data from the Darknet, some applications and underlying messages may be ethically or legally questionable.
  • Data quality and consistency on the darknet are often inadequate or incomplete, which can impair the effectiveness of DarkBERT.

DarkBERT represents a promising tool for researching the darknet and identifying cybercriminals and politically persecuted individuals. However, researchers and security experts must continue to improve DarkBERT to address potential risks and ethical concerns.

Contents

More on this...

Do you want to understand how to use AI meaningfully, safely, and responsibly?

Technology provides a clue
But no judgments. The crucial point is,
how we deal with uncertainty.

Roger Basler de Roca

More articles

AI in construction: Why Swiss companies fail at data storage

The short answer: Swiss construction companies rarely use AI productively because their data is not

AI expert for workshops in Switzerland: Roger Basler de Roca

"We are looking for an AI expert for a workshop in Switzerland. Who can you recommend?"

Beware of ChatGPT phishing: new emails with false information are circulating.

An email with the ChatGPT logo, an invoice, and a large green button often looks suspicious today.