24 November 2024 by Mondo2.0
The mysterious disappearance of web pages
I am taking my cue from the Pew Research Center study entitled “When Online Content Disappears” to share some sociological reflections on the Internet. The opening statement in the study is clear and unequivocal:
“A quarter of the web pages that were published between 2013 and 2023 are no longer accessible”
As often happens on the Internet, a behavior, a dynamic, occurs regardless of the place, the content or the author.
This phenomenon, called “digital decay”, affects blogs, occasional websites, government websites and Wikipedia pages alike. It opens the way to the “Digital dark age”, namely the impossibility of retrieving digital material published in our era because it is stored in obsolete cloud systems and hardware, or hosted in temporary spaces and, in any case, no longer available.
The percentages reported in the study give us food for thought:
- 21% of web pages on government websites contain at least one broken link
- 54% of Wikipedia pages contain at least one link in the “References” section that points to a page that no longer exists. More precisely, 11% of all links on Wikipedia are no longer accessible. On about 2% of the source pages containing reference links, every link on the page was broken or otherwise inaccessible, while another 53% of pages contained at least one broken link.
To this is added the subsequent analysis carried out in the same study, namely the analysis of a sample of tweets published in the spring of 2023 and followed over a period of three months:
- Nearly one in five tweets is no longer publicly visible on the site just a few months after publication. In 60% of these cases, the account that originally posted the tweet was made private, suspended or deleted altogether.
- Some types of tweets tend to disappear more often than others. More than 40% of tweets written in Turkish or Arabic are no longer visible on the site within three months of publication.
After this series of figures, the alarmist part should begin, the one that would allow me to significantly increase the number of readers of this blog, with statements such as “web pages are disappearing from the Internet!”, “the Internet is disappearing”, “Mystery!”, “Conspiracy!”, “All the Internet output of the 21st century is disappearing!”.
The numbers do not lie: the pages and tweets are no longer there, a large part of the digital material created in this quarter-century will gradually be lost, but the picture is certainly more complex.
Let us take a step back and look at the other media:
- Anything said on the radio, given the nature of the medium, is valid only at the moment of the announcement, for a few moments. It may be picked up by other media or temporarily recorded in a database, but by its nature it is instantaneous, volatile.
- Newspapers, by their very origin and name, have daily value and attention. Newspaper archives that preserve copies of newspapers are quite rare.
- Television news, talk shows and in-depth science programs can be viewed again in large online databases; they are often available for a limited period of time, at most a few months (also for obvious reasons of interest and the cost of cloud storage).
Conversely, historical digital databases are available over the long term only for content that we can define as “excellent”, sometimes exclusively for professionals.
So why should content published on the Internet last more than ten years? Why are we surprised that a significant part of what is published on the Internet disappears?
Probably many of us think of the Internet as a container, a digital display case, an extremely personal digital space in which to leave an indelible mark, but that is not the case.
The primary purpose of the Internet is not the preservation of knowledge; it is to act as a sounding board for content and emotions, it is the (temporary) dissemination of information.
In practice, we are confusing a large drum, through which all of us can make a great deal of noise, with a large library where we can preserve books.
Digital repositories are quite another matter, for example in academia, where they are useful for preserving and disseminating knowledge, but, as mentioned, these are “excellent” contents, not just any web page or tweet.
Web pages published on the Internet naturally have a short life; digital decay is a normal, physiological phenomenon, also given the immense amount of content produced online every day. Web pages on institutional sites or referenced by Wikipedia are no exception.
Over a decade, search engines change, as do the ways in which they give attention to information; communication dynamics change; public bodies and private companies change their names; organizational charts and organizations change; methods of use evolve and, consequently, so do the strategies for designing and defining menus on websites, where page URLs (addresses) are located. Posts and tweets even more so: they are born in the moment and disappear in a hurry; they last as long as a chirp.
Looking at the glass as half empty, or rather noticing a few cracks in the glass, it is also clear that the extreme dynamism that characterizes the Internet encourages the spread of false news.
The inability to trace the source, the origin of the news, allows anyone, for a few moments or a few days, to say anything, open debates that polarize opinions, and provide statistics with no supporting evidence.
The ability to open and close temporary accounts supports the publication of partially or completely false news.
The use of artificial intelligence to generate new content will increase the online presence of temporary, occasional and, unfortunately, partly false content.
It will be possible, indeed it is already possible, to capture the attention of consumers and voters by generating web trends based on false news and then, with a quick sleight of hand, make the table and cards, the account and the news disappear.
The antidote to all this is not the creation of definitive content on the Internet, nor chasing temporary accounts and content, which is not possible, but rather developing in the new generations the critical ability to analyze information and sources.
Who is telling me what? Why are they saying it?
Who knows, perhaps faced with “A.I. dark” services that generate false news, we may end up using “super A.I.” services that verify the reliability of the source. Who knows.
While waiting to better understand what lies ahead, I will conclude with a paradox and two references to old articles:
- The website (of an extremely reliable Italian newspaper) where I originally found the references to the study cited at the beginning of the article contains, on the very same page, links that no longer work. A great demonstration of consistency.
- The Mondoduepuntozero Blog has been active for more than 11 years; our very first article, dated March 19, 2013, is still available. More than 4,000 days have passed; the article was dedicated to Google, proving that not everything passes quickly on the Internet: we at Mondoduepuntozero are still here.
- Our January 1, 2017 article, “Post Truth: Hoaxes on the Internet”, is still relevant and, in addition to containing a few suggestions, ends with the following statement:
“The Internet, as we often say, is neither the best nor the worst of all possible worlds; it is simply a mirror of our time.”
Mondoduepuntozero
