Wednesday, 26 July 2023

Trends in Web3 – Part 1: A look at blockchain domains

by David Barnett, Rebecca Newman, Tom Ambridge and Richard Ferguson

Introduction

The world of Web3 - the so-called 'third generation' of web content which has emerged over the last few years in response to a series of technological developments - continues to generate significant amounts of discussion and speculation. A number of the associated applications present the potential for new types of infringements, brand abuse and online risks, and warrant close attention by brand owners and general Internet users alike.

In this article we consider trends in the Web3 landscape - with a particular focus on blockchain domains - and how these may be driven by external factors, such as developments in artificial intelligence (AI). We also discuss the implications for brand owners and the associated brand protection considerations.

Overview of Web3 concepts

i. Technical definitions

Web3 (a.k.a. 'Web 3.0') is a general term referring to decentralised content on the Internet (i.e. organised on a peer-to-peer basis and without reliance on authoritative hosting providers), with a particular focus on blockchain technologies. The term 'Web3' reflects the emergence of new technologies which provide immersive experiences for users and encourage freedom of speech. It follows on from the definition of Web2 in the early-2000s, to indicate a transition from the 'read-only' days of the early Internet into an ecosystem more dominated by user-generated content.

A blockchain is a publicly accessible digital ledger in which transactions are recorded. It is cryptographically sealed and cannot be modified after its contents are recorded. Blockchains form the basis of many digital currencies (such as Bitcoin), but also have a number of other applications, such as supply-chain control by brand owners. In terms of brand protection considerations, two related concepts are of particular relevance:

  • NFTs (non-fungible tokens)[1] - NFTs are cryptographic collectibles comprising any of several types of asset or media. Any digital file can be converted into an NFT through a process known as minting, whereby ownership is recorded on a blockchain. NFTs can take a number of different forms, but are most commonly associated with graphics files (e.g. artworks, branded imagery, etc.) and other types of digital content (such as audio or music files). Brand owners are increasingly incorporating NFTs into their business models, such as the production and trade of virtual branded items (e.g. items to be worn by avatars in virtual-reality environments - part of the 'metaverse', the name given to a generalised connected environment of 3D virtual worlds.
  • Blockchain domains - Like regular domains, blockchain domains consist of a second-level domain name and an extension (with specific examples including .eth, .crypto, and .bit), and can be used in a number of different ways, including the construction of decentralised websites (which have special access requirements, such as the use of a dedicated browser like Brave, or a browser plug-in), as memorable wallet addresses for sending and receiving cryptocurrency, or as hosting infrastructure for programs to be run as apps. Blockchain domains are recorded, together with their ownership details, on a blockchain (i.e. are not hosted on a server, or recorded in a regular registry zone file) and, unlike regular domain names, are not governed or regulated by ICANN (the Internet Corporation for Assigned Names and Numbers). They are offered by specialist providers and, although the costs may be higher than for traditional domains, are in many cases offered for registration for a longer period than gTLDs, or involve only a one-off cost to own for ever.

Blockchain domains are attractive to many users because of the inherent security associated with the blockchain infrastructure, and their resistance against traditional blocking or censorship methods. For those with interests in other areas of the Web3 ecosystem, use of blockchain domains may be a natural choice.

Future development of native support of blockchain domains by mainstream web browsers is also likely to significantly drive increased adoption by users. Already, blockchain domain operators are seeking technical workarounds to drive the interoperability of blockchain domains with regular browsers. One example is eth.link, a service allowing .eth blockchain domains to be accessed via DNS, by appending '.link' to the blockchain domain name[2].

ii. Brand protection implications

From a brand monitoring point of view, blockchain domains are generally difficult to identify, both because of the absence of zone files (which, for regular domains, provide comprehensive lists of registered domains across the individual domain extensions, or TLDs (top-level domains)), and because of the specific website access requirements. One commonly-used technique to circumvent this difficulty can be to search for references to the blockchain domain names being traded in NFT marketplaces (for example, where the current owner can offer the sale of a domain to another interested party, which may be a brand owner or a would-be infringer), and some blockchain domain providers also provide searchable databases of registered domains. However, more robust methods are likely to require direct monitoring of the content of the blockchains themselves, or searches across databases of transactions which have occurred on a specific blockchain.

Enforcement options are also currently limited, with one option being just to take down infringing listings offering the sale of a blockchain domain from the Web3 marketplace - although this does not deactivate the domain name itself or change its ownership. Some blockchain domain providers are becoming more mindful of the risks posed by cybersquatters[3], and offer brand owners the ability to block third-party registrations (similar to the Trademark Clearinghouse (TMCH) programme for new gTLDs) or to claim ownership of trademarked names. However, these blocks are at the discretion of the domain providers, making them subject to change and in need of periodic monitoring.

Additionally, wallet addresses associated with specific blockchain domain or NFT owners - as might be available through public records, Web3 marketplaces, or blockchain domain providers - can be used as the basis of an investigation to identify additional associated information relating to the entity in question. In some cases, it may also be possible to submit a court order to the service provider for the disclosure of collected data.

As a further brand protection initiative, brand owners may also wish to consider proactively defensively registering key domain name keyword strings across relevant extensions. This approach is generally more cost-effective than attempting to subsequently acquire domain names of interest.

The changing landscape

A number of recent factors - primarily driven by developments in AI technologies - may very well impact on the role of Web3 within the wider Internet landscape, even if (as we shall see) we have not yet seen any major new growth in the level of uptake of the relevant technologies.

The developers of AI products and services such as ChatGPT are increasingly using the data held by platforms such as Reddit and Twitter as an input for their training models, resulting in the introduction of initiatives by these platforms to prevent or monetise this activity. In April 2023, Reddit announced plans to start charging for use of its data API[4], resulting in a number of subreddits (communities) - or associated applications - being made private, or shutting down altogether[5], ('going dark') in protest[6,7]. In July, Twitter introduced a measure to limit the number of posts a user could read per day[8], to prevent "extreme levels of data scraping and system manipulation"[9]. The actions taken by online platforms to remove and restrict access to content are widely seen as being contrary to a fundamental characteristic of Web2, namely the ability of individual users to create and curate their own content, and choose which content to consume from other users.

It is interesting to note that Web3 and the metaverse has recently been declared 'dead' by some commentators[10], following moves by Meta and Mark Zuckerberg - together with other industry leaders - away from these areas, in favour of development of generative AI[11]. However, it is perhaps that very trend - and the associated reactions by service providers - which may point us back to the need for a Web3, being inherently decentralised and less prone to regulation and restriction. Could these technologies see a new lease of life as a safe harbour for the community-owned content which used to be the province of Web2[12]? The answer is a resounding 'yes', according to AI commentator Alex Valaitis, who recently tweeted that "AI becomes stronger the more centralized it gets, which is exactly why we need Web3 as a counterweight[;] think of public blockchains as the last bastion of the open internet"[13].

The existence of the Web3 Domain Alliance[14] is also noteworthy. It features many of the major Web3 service providers as members, and is intended to drive "consumer protection, preventing naming collisions, fair and open use of intellectual property in the industry, and interoperability of blockchain naming systems"[15]. As part of this initiative, Unstoppable Domains announced in February 2023 that it would not enforce a key patent - relating to the use of smart contracts[16] in the blockchain domain naming process - against other members of the group[17].

The Web3 ecosystem also presents its own problems, however, such as the ongoing US litigation between Unstoppable Domains and Wallet Inc. over rights to offer domains across the .wallet extension[18].

Observed trends in blockchain domains and Web3 content

There is evidence of a significant amount of (at the very least, legacy) interest in Web3 concepts; searches across the (gTLD) zone files provided by ICANN show that, as of July 2023, there are over 300,000 registered (regular) domain names containing the Web3-related keywords 'web3', 'nft', 'blockchain' or 'metaverse'.

Currently, there are around 7 million blockchain domains registered, with two of the most popular providers - Ethereum Name Service ('ENS') (which offers .eth domains on the blockchain associated with the Ethereum cryptocurrency) and Unstoppable Domains (offering blockchain domains across a range of more than ten extensions) - having provided 2.7 million[19] and 3.6 million[20] registrations, respectively.

In the remainder of this article, we consider activity surrounding .eth registrations as a proxy for the overall blockchain domain landscape, both because of the popularity of the Ethereum Name Service and because of the ready availability of associated statistics available through information and tools provided by Dune Analytics.

Dune states that, as of 17 July 2023, there are 2,719,569 active ENS (.eth) blockchain domains[21]. The registration history of these domains is available back to May 2019, and is shown in Figure 1.

Figure 1: Monthly numbers of .eth blockchain domain registrations

The statistics show a very large peak in activity covering roughly the calendar year of 2022, after which levels of registrations appear to have dropped off for now. This is consistent with the flurry of activity where available three- and four-digit domains were being rapidly registered, with the monthly trading volume peaking at $44.3 million in May 2022[22].

It is also possible to extract more granular data, looking at the individual blockchain domain registrations (names and registration dates), for the last year (July 2022 – July 2023)[23] (Figure 2).

Figure 2: Daily numbers of .eth blockchain domain registrations (July 2022 – July 2023)

Within this dataset of 1.47 million blockchain domains, we consider the prevalence of domains with names containing each of the top ten most valuable global brands in 2023 (according to data provided in the latest Kantar BrandZ study[24]). This dataset (see Table 1 and Figure 3) provides a measure of the likely level of potential brand infringement across the blockchain domain landscape.

Brand string
                                       
No. registered .eth domains
(July 2022 - July 2023)
                                                   
  apple 902
  google 625
  microsoft 249
  amazon 926
  mcdonalds 136
  visa 301
  tencent 43
  vuitton 119
  mastercard 66
  coca(-)cola * 160

* The hyphen in the brand string is optional, so the data considers examples containing 'cocacola' or 'coca-cola'.

Table 1: Total numbers of .eth blockchain domains registered between July 2022 and July 2023 with names containing each of the top ten most valuable global brands in 2023

Figure 3: Monthly numbers of .eth blockchain domain registrations with names containing each of the top ten most valuable global brands in 2023

It is also worth noting that the inclusion of Unicode support in the blockchain domain infrastructure allows special characters such as emojis to be included in the domain names[25]. Consequently, the dataset includes examples such as those shown in Figure 4, many of which have the potential to be used to create highly deceptive and/or purportedly official websites.

Figure 4: Examples of branded .eth blockchain domain names with names including special characters

A similar previous study published on the DNS Research Federation blog has highlighted the potential for such domains to be used fraudulently or for other infringing purposes, and identified multiple instances of prolific serial registrants, each in ownership of over 100 branded domain names[26]. The difficulties with monitoring and enforcement across the blockchain landscape also makes these domains attractive to cybersquatters[27].

Currently, very few of the branded blockchain domains in the dataset resolves to any significant content, although a small number of live sites or active website responses were identified (Figure 5). These include one webpage (Figure 5(iii)) analogous to a server index page, which is sometimes seen with sites under development as a precursor to subsequent, more significant website content, or when 'hidden' (potentially harmful) content is present in one of the subdirectories.

(i)

(ii)

(iii)

(iv)

Figure 5: Examples of live websites or other active webpage responses associated with .eth blockchain domains with names containing any of the top ten most valuable global brands in 2023 - (i) googleisadog.eth; (ii) whalevisa.eth (potentially unrelated to the Visa brand); (iii) tencentglobal.eth; (iv) googlenoodle.eth

Conclusions

These observations raise a number of questions about the likely direction of Web3 trends going forward. Arguably, there is a case to be made that the peak in activity in blockchain domain registrations took place in 2022, and has now greatly subsided. However, it may simply be that this time period simply represented a 'golden age' for registrations, when significant numbers of the highly-desirable available domain names were snapped up by prospectors; the creation of a pre-existing landscape to which future activity will be added. It is also possible that the nascent nature of Web3, and uncertainty by brand owners over which department should take responsibility for blockchain domains, has resulted in reluctance by corporations to embrace and adopt the technologies.

Furthermore, there must also be questions surrounding the use of an analysis of domains with just a single extension, on a single blockchain, as a proxy for the whole Web3 ecosystem.

Overall, however, it seems reasonable to assert that AI technologies - and other technological and societal developments - may well drive a resurgence of interest in Web 3. Given the scale of legacy activity in the associated areas, and the numbers and nature of associated infringements, it seems advisable for brand owners to be mindful of blockchain domains targeting their brands, and to carefully consider their own brand-protection strategies in the Web3 arena.

References

[1] https://www.linkedin.com/pulse/rise-nft-david-barnett

[2] https://eth.link/

[3] https://www.brandsec.com.au/blockchain-domains-and-cybersquatting/

[4] https://techcrunch.com/2023/04/18/reddit-will-begin-charging-for-access-to-its-api/

[5] https://www.reddit.com/r/apolloapp/comments/144f6xm/apollo_will_close_down_on_june_30th_reddits/

[6] https://news.sky.com/story/reddit-blackout-thousands-of-communities-are-doing-dark-today-heres-why-12899280

[7] https://www.theverge.com/2023/6/30/23779519/reddit-third-party-app-shut-down-apollo-sync-baconreader-api-protest

[8] https://www.bbc.co.uk/news/technology-66093324

[9] https://twitter.com/elonmusk/status/1675187969420828672

[10] https://www.splunk.com/en_us/blog/learn/blockchain-web3-dead.html

[11] https://www.businessinsider.com/metaverse-dead-obituary-facebook-mark-zuckerberg-tech-fad-ai-chatgpt-2023-5

[12] https://cointelegraph.com/news/how-adoption-of-a-decentralized-internet-can-improve-digital-ownership

[13] https://twitter.com/alex_valaitis/status/1674840503861248018

[14] https://www.web3domainalliance.com/

[15] https://cointelegraph.com/news/web3-domain-alliance-expands-with-51-new-members

[16] https://www.investopedia.com/terms/s/smart-contracts.asp

[17] https://fortune.com/crypto/2023/02/22/unstoppable-pledges-patent-non-aggression-pact-across-expanded-web3-domain-alliance/

[18] Unstoppable Domains, Inc. v. Wallet Inc. et al. 1:2022cv01231

[19] https://ens.domains/

[20] https://unstoppabledomains.com/

[21] https://dune.com/makoto/ens

[22] https://dappradar.com/blog/best-blockchain-web3-domain-names-services

[23] https://dune.com/makoto/ens-released-to-be-released-names

[24] https://www.kantar.com/inspiration/brands/revealed-the-worlds-most-valuable-brands-of-2023

[25] https://nptacek.medium.com/experimenting-with-ens-c88bfe7ed246

[26] https://dnsrf.org/blog/brand-names-in-blockchain-domains---new-frontier-for-brand-owners/index.html

[27] https://www.thefashionlaw.com/the-rise-in-blockchain-domains-presents-risks-opportunities-for-brands/

This article was first published on 26 July 2023 at:

https://www.iamstobbs.com/opinion/trends-in-web3-part-1-a-look-at-blockchain-domains

Monday, 3 July 2023

An overview of the concept and use of domain-name entropy

Introduction

In this article, I present an overview of a series of 'proof-of-concept' studies looking at the application of domain-name entropy as a means of clustering together related domain registrations, and serving as an input into potential metrics to determine the likely level of threat which may be posed by a domain.

In our previous studies, we utilised the mathematical concept of Shannon entropy[1], providing a measure of the amount of information stored in a string of characters (or, equivalently, the number of bits required to optimally encode the string). The idea was applied to the second-level domain name (SLD) part of each domain (i.e. the portion of the domain name before the dot - such as 'google' in 'google.com'), and broadly means that short domain names, or those with large numbers of repeated characters, will have low entropy values, whereas longer domain names, or those with large numbers of distinct characters, will have higher entropy.

The background to this analysis is the fact that domains registered for egregious purposes (such as spamming, malware distribution, or botnet creation) may be more likely to be registered in bulk by bad actors using automated algorithms[2], which typically results in the generation of long, non-sensical (i.e. high entropy) domain names, which have the added benefit of not containing brand-related keywords and are typically therefore harder to detect using classic brand-monitoring techniques. The idea is that domains registered by a particular infringer for a specific campaign are likely all to be generated using the same algorithm, and may therefore have similar or identical entropy values.

Overview of previous studies

In our initial proof of concept[3], we considered the set of all domains registered on a particular day - a sample of around 205,000 domains. The advantage also of considering a set of domains with a common registration date is that it presents the possibility for one or more groups of automated bulk registrations (which are typically all registered at the same time) to be present.

Within the dataset, a range of domain entropy values was present, from a minimum of 0.000, to a maximum of 4.700, and with 92.3% of the dataset having values below 3.500. (see Figure 1). The top 1,000 highest-entropy domains (i.e. the top 0.49%) had entropy values in excess of 3.823, and accounted for the majority of examples which appeared visually to feature 'random' SLD strings. Within this high-entropy subset, a number of additional characteristics were indicative that many may have been registered for nefarious purposes, including the prominence of use of consumer-grade registrars and privacy-protection services, and the extent of the presence of active MX records amongst these new registrations (in 27.5% of the cases - indicating that these domains have been configured to be able to send and receive e-mails and therefore could potentially be associated with phishing activity).

Figure 1: Cumulative proportion of domains with entropy less than the value shown on the horizontal axis, from the dataset in the initial proof-of-concept study

Indeed, at least one apparent 'cluster' of suspicious registrations was found to be present within the dataset, comprising a group of 125 .buzz ('dot-buzz') domains, all with an identical high entropy value (3.907), registered via a common registrar and associated with groups of similar IP addresses. At the time of analysis, many of the domains registered to Chinese-language, gambling-related websites, likely representing either an affiliate revenue generation scheme, or 'dummy' content serving to 'mask' higher-threat content which may only have been visible in specific geographic regions, or which may have been planned for subsequent upload.

In a follow-up study[4], I considered a month's worth of registrations of domains with names containing any of the top ten most valuable brands in 2022. Similarly, the high entropy domain names within this dataset included groups of apparently related, coordinated 'clusters' of domains, several of which appeared intended for fraudulent use and were consistent with registration via automated generation algorithms. For example, seven of the top eight domains in the dataset (by entropy values) had similar names of the form 'google-site-verificationXXXXXX.com' (or .net) (where 'XXXXXX' was a long string of apparently random characters), and a series of groups of 'microsoft' examples was identified, including keywords such as 'cloudworkflow', 'netsuites' and 'cloudroam'.

Comparison with other work

Other studies taking similar approaches to the analysis of domain entropy also reach similar conclusions. For example, an analysis outlined in a blog posting by Tiberium[5] states that the use of an entropy threshold of >3.1 (as an indicator of potential concern) correctly classifies 80% of NCSC malicious domains, and incorrectly classifies only 8% of the top 1000 most popular (legitimate!) domains overall (cf. Table 1).

Domain name
                                       
Entropy value
                           
  google.com 1.918
  youtube.com 2.522
  facebook.com 2.750
  twitter.com 2.128
  instagram.com 2.948
  baidu.com 2.322
  wikipedia.org 2.642
  yandex.ru 2.585
  yahoo.com 1.922
  whatsapp.com 2.500

Table 1: Entropy values of the SLDs of the top ten most popular websites according to Similarweb[6]

Additionally, an article published by Splunk[7] looking at the entropy values of fully qualified domain names, i.e. also including subdomain names - also states that high-entropy examples are consistent with the use of domain generation algorithms, and may be indicative of association with malware (e.g. in 'beaconing') and other web exploits. Comparable approaches and conclusions can also be found in a range of other studies[8,9,10], with some finding improvements in the reliability of threat determination through the use of alternative measures such as relative entropy (essentially, a comparison against the character distribution observed in a dataset of known legitimate domains, so as to provide a better measure of the randomness arising from automated algorithmic registrations)[11].

Conclusions

Domain-name entropy analysis has applications in at least two key areas of brand protection. The first of these is the ability to 'cluster' together related infringements, which has a number of benefits, including the ability to identify serial infringers and instances of bad-faith activity, for targeted and effective bulk enforcement actions. The second key area is as an input into algorithms to quantify the likely level of threat which may be posed by an online feature such as a new domain registration. Threat determination is essential in allowing prioritisation of results for analysis, enforcement, or content-change tracking.

All other factors being equal, there is some indication that high-threat domains - particularly those associated with automated registrations by domain-name generation algorithms - may have a tendency to sit at the higher-entropy end of the spectrum (and, furthermore, that domain names generated using a particular algorithm may be likely to have similar entropy values). This statement runs alongside the assertion that legitimate domains may (in general) be more likely to have lower entropy values, particularly where there is a desire for legitimate businesses to utilise strongly branded, short, memorable web addresses - as can be seen in many of the globally most popular websites.

References

[1] https://arxiv.org/ftp/arxiv/papers/1405/1405.2061.pdf

[2] https://interisle.net/sub/CriminalDomainAbuse.pdf

[3] https://www.linkedin.com/pulse/investigating-use-domain-name-entropy-clustering-results-barnett/

[4] https://www.linkedin.com/pulse/entropy-analysis-registered-domain-names-relating-top-david-barnett/

[5] https://www.tiberium.io/blog/chapter-2-classifying-domains-through-string-entropy/

[6] https://www.similarweb.com/top-websites/

[7] https://www.splunk.com/en_us/blog/security/random-words-on-entropy-and-dns.html

[8] https://hurricanelabs.com/blog/dns-entropy-hunting-and-you/

[9] https://www.logpoint.com/en/blog/embracing-randomness-to-detect-threats-through-entropy/

[10] https://suleman-qutb.medium.com/use-of-shannon-entropy-estimation-for-dga-detection-9ded275795ca

[11] https://redcanary.com/blog/threat-hunting-entropy/

This article was first published on 3 July 2023 at:

https://circleid.com/posts/20230703-an-overview-of-the-concept-and-use-of-domain-name-entropy

Thursday, 25 May 2023

The 'Millennium Problems' in Brand Protection

As the brand protection industry approaches a quarter of a century in age, following the founding of pioneers Envisional[1] and MarkMonitor[2] in 1999, I present an overview of some of the main outstanding issues which are frequently unaddressed or are generally only partially solved by brand protection service providers. I term these the 'Millennium Problems' in reference to the set of unsolved mathematical problems published in 2000 by the Clay Mathematics Institute[3], and for which significant prizes were offered for solutions. Like their mathematical counterparts, the unsolved problems in brand protection will present significant benefits for any service providers able to develop and offer comprehensive solutions.

Brand protection basics

In their most basic sense, brand protection solutions generally consist of two components: monitoring (or, strictly, detection) of brand-related content on the Internet, and enforcement action to achieve the removal of infringing material. Monitoring is most usually carried out using technological solutions intended to identify relevant material on the Internet, across a range of relevant channels, typically using a combination of methodologies, namely: (i) Internet metasearching (i.e. the submission of relevant query terms to search engines) and web crawling; (ii) analysis of domain-name zone files (see Problem 2), to identify domains with names including brand-related terms (or variants); (iii) direct monitoring / searching on known sites of interest (see Problem 1); and (iv) other techniques, such as the use of spam traps and webserver logs, as used in phishing detection technologies[4]. Many service providers will also make use of automated analysis tools, which can inspect the content of the identified webpages, and categorise and prioritise these results accordingly.

The 'Millennium Problems'

1. Social media monitoring

Whilst monitoring of content across social media platforms is a well-established element of many brand-protection service providers' product suites, it frequently remains extremely difficult to achieve anything approaching a comprehensive level of coverage. There are a number of reasons why this is the case. In general, social media content is most usually addressed using the 'direct site searching' approach (that is, using the search functionality typically in-built to the platforms themselves as a means of returning results), though some providers also have access to direct data feeds from the platforms (e.g. through an API). In general, a variety of types of content may be of interest, including brand references in usernames (e.g. associated with fake profiles), and the content of postings (e.g. associated with fraud, the sale of counterfeits, the spread of malware, brand disparagement, etc.) and elsewhere (including imagery, sponsored advertisements, and so on).

The main difficulty with the 'direct search' approach is that results presented to a user are often limited (sometimes significantly) unless the user is logged in to the social media platform. This can be circumvented by configuring a brand-protection monitoring tool to present itself to the platform as if it is a real user (with a registered account, handle (username) and password), or simply through the use of manual searches. Both of these approaches typically require the use of 'dummy' accounts and may be in contravention of the terms and conditions of the platforms themselves.

Other technological issues may also be problematic. Many social media platforms return results on an 'infinite scroll' basis (where additional results are continually added to the webpage as the user continues to scroll down through them), often with no indication of the total numbers of results which may be present, and many platforms also have specific access requirements, such as functionality only to be accessed via a mobile app (see Problem 7). Similarly, monitoring can be further complicated by sites where content is protected via the requirement to enter a CAPTCHA code, for example. It is also typically the case that the exact results returned to a user will be highly personalised, and dependent on their browsing history, interests, location, and personal demographic.

Some of these issues can be addressed through the development of partner relationships by brand-protection service providers with the platforms themselves. However, even in cases where the platforms are amenable to this approach, some of the above technological issues may remain difficult to address.

2. Comprehensive ccTLD monitoring

Another of the core elements of many brand protection service offerings is often a domain monitoring capability; that is, the ability to identify domains whose names include the name of the brand being infringed (and/or other relevant keywords). As a special subset of general Internet content, branded domain names are often of particular interest by virtue of their greater visibility (e.g. higher ranking in search-engine results) and the more explicit nature of the IP abuse (and an associated greater range of enforcement options)[5]. Branded domain names have been noted in many previous studies as being popular with bad actors in the creation of infringing content of a variety of types, including phishing sites[6], sites offering the sale of counterfeits, and sites claiming false affiliation or including disparaging content.

The primary source of data for domain monitoring is usually the analysis of zone files, which are data files published by the registry organisations responsible for overseeing the infrastructure of each individual TLD (top-level domain, or domain extension - such as .com), and which contain a list of all existing registered domains across that extension. By comparing the content of a zone file with that from the previous day, it is possible to identify new domain registrations (as well as dropped, or lapsed, domains) and filter this list for those examples containing a brand name or keyword of interest. Domain monitoring solutions can (and, in general, should) also make use of zone-file analysis to allow identification of the full pre-existing 'landscape' of registered domain names of interest, across the TLDs in question, at the commencement of monitoring (so-called 'baseline' analysis). The most sophisticated domain monitoring solutions can also automatically check for variations of the brand strings (such as typos), which are frequently used by infringers to construct deliberately deceptive domain names[7,8].

Zone files are generally available for most gTLDs (generic, or global, TLDs such as .com, .net, etc.) plus the new-gTLDs which have been launched in the period since 2012[9], but are often not published (or may not be comprehensive) by the registry organisations responsible for other TLDs, particularly the country-specific examples (ccTLDs). For this reason, detection of relevant domains across ccTLD extensions is typically incomplete, and a number of techniques may typically be used in order to fill in the gaps. These might include parallel look-ups (checks for domains with the same second-level domain name - i.e. the part of the domain name to the left of the dot - as examples identified through zone-file analysis), exact-match queries (regular searches for the existence of domains with second-level domain name strings of particular relevance, such as a brand name), and Internet metasearching. However, each of these approaches has its own limitations and, even when all taken together, there can always be domain names of potential concern which are not detected through any of these methods. The next generation of domain monitoring solutions will need to better address these shortcomings, potentially involving factors such as the use of improved algorithms to 'guess' candidate domain names for checking, and/or the use of more comprehensive indexes of Internet content. Additionally, the building of specific relationships with country registries - potentially combined with regulatory changes regarding the availability of zone files - may also be relevant.

3. Third-party subdomain monitoring

The subdomain is the section of a URL prior to the domain name, from which it is separated by a dot (e.g. 'translate' in 'translate.google.com'). The owner of a domain name can create whatever subdomains they wish, and can point these URLs to associated web content (via the configuration of DNS settings). Accordingly, subdomains can be used to create brand-related URLs, and can be associated with many of the same types of infringements as domain names themselves[10]. Subdomain-based abuse can also be particularly attractive to infringers, both because it avoids the requirement to register a brand-specific domain name[11] (which bad actors know can easily be detected by brand owners employing domain-monitoring services) and because there can be a low cost associated with the creation of the URL, particularly where a service provider allowing the free registration of personalised subdomains (such as blogspot.com) is used.

Consequently, the ability to monitor generally for brand references in the subdomain name of arbitrary URLs can be of great value. Note that this is distinct from the (relatively much simpler) problem of monitoring the existence and content of subdomains of official domains under the ownership of the brand owner 'internal' subdomain monitoring), since all of the relevant information is contained in the DNS configuration files held by the brand owner's domain-name management service provider.

Conversely, the identification of brand-related subdomains on third-party ('external') domain names is much more difficult. In many cases, this is achieved purely using Internet metasearching techniques (i.e. finding only content which is indexed by search engines in response to brand-specific query terms). Whilst this does mimic the search techniques used by general Internet users (and thereby identify the 'highest-visibility' content), it will in general not find all potentially threatening content (e.g. URLs to which traffic is driven through other means, such as links in spam e-mails). This problem can be mediated to some degree through the use of other techniques, such as passive DNS analysis or certificate transparency (CT) analysis, or via explicit queries for the existence of specific subdomain names of interest. However, these techniques require prior identification of the specific domains to be monitored; generalised identification of brand-related subdomains remains a much harder problem to solve.

4. Circumventing site blocking and geoblocking

Site blocking and geoblocking are two long-established problems in brand monitoring. The former arises when a monitored site becomes aware of repeated search queries from a particular source, and restricts access to the site from the IP address in question. A site owner may choose to do this for a number of reasons, including protection of website performance (e.g. in preventing DDoS attacks), or for compliance with their own terms and conditions (e.g. where they state that information is not to be collected for commercial purposes, such as by brand-protection service providers). Geoblocking (or geotargeting) is a related issue, whereby the visible content of a website may vary depending on the geographical location of the visitor. Again, this may be implemented by a site owner for a range of reasons, including the tailoring of content to a local audience, search-engine optimisation, security, or legal compliance[12]. However, geoblocking can also be employed by infringers as a means of evading detection, and can also present difficulties in enforcement, where it may be necessary to demonstrate exactly what content is visible from a specific remote location.

The solutions to these issues, from a brand-protection point of view, are relatively simple in principle, generally involving the use of proxies (standalone external machines serving as intermediate 'hops' through which search queries from a brand-protection service provider are routed, so as to 'mask' the originating IP address) in a range of remote locations, and/or (particularly for site blocking) the building of relationships with the sites being monitored, so that the monitoring service provider can gain permission for collecting the data. However, in practice this requires a great deal of investment in building the required infrastructure (such as hosting and maintaining the necessary proxies, and configuring the monitoring software to communicate with them) and establishing the necessary relationships. Furthermore, the construction of appropriate user interfaces to visualise and interpret the relevant information (such as the ability to compare the content of a particular website across a range of different user (i.e. proxy) locations, in cases where geoblocking or geotargeting may be an issue) can also be a complex prospect.

5. Clustering and open-source intelligence analysis

The subject areas of clustering and open-source intelligence (OSINT) are generally of greatest relevance for entity investigations, i.e. the process of using Internet searches to build a portfolio of information relating to an identified individual or website of interest. Such information can be used for a range of purposes, including background for on-the-ground investigations or goods seizures, or for legal cases, but can also be useful background for enforcement actions (e.g. in identifying clusters of related infringements for efficient bulk takedowns in a single action).

A number of technological solutions exist for visualising the links behind related entities, on the basis of common shared characteristics (such as e-mail addresses, telephone numbers, web-hosting information such as IP addresses, and so on) - i.e. 'clustering', but it is often the case that the characteristics themselves require identification through manual analysis processes. A great deal of additional efficiency can be built into the process, however, through the use of monitoring and analysis tools which can identify and extract this information automatically. This is relatively more straightforward in cases where the data can be extracted in a consistent manner (e.g. performing an IP-address look-up for any identified website of interest), and/or where the information is contained in a known location on a webpage with a fixed, pre-defined format (the 'contact details' section of a social-media profile page), such that a web scraper can be configured to pull out the content. It is a considerably more difficult enterprise to extract such information from general webpages where the structure of each page is not known in advance. In these cases, the approach generally needs to be based on the configuration of monitoring tools which are able to extract text-strings with the general format of (say) an e-mail address or telephone number. This then typically requires an element of post-processing to 'clean' and standardise the data. The next generation of clustering tools are likely to make extensive use of artificial intelligence in order to do this, in addition to also then drawing out insights between the clusters thus produced.

6. Dark Web monitoring

Dark Web content is the general name given to online material for which there are special access requirements; however in the context of online brand monitoring, it is usually taken to refer to content which is only accessible via the Tor network (a decentralised network involving the use of encrypted communications, and connections via multiple hops between Tor servers (proxies) - also known as relays or nodes). The Tor network - which is accessed using specially enabled browsers - can be used to view regular ('surface web') Internet content (and is one option open to users for whom anonymity is important), but is more usually used to access websites with the .onion extension, i.e. those which are only accessible from within the network[13].

The Tor network of .onion websites includes a range of different content types, but is notorious for illegal and infringing content and, as such, can be a key area of interest for brand monitoring. However, many brand protection service providers offer only limited capabilities in this area. This is for a number of different reasons. One significant factor is that the Dark Web is essentially unregulated, frequently with no available links to 'real-world' contact details, and extremely limited enforcement options against infringing content. However, even in cases where takedown is not possible, intelligence on the content can be extremely valuable - one example may be on 'carder' websites, on which stolen financial credentials are traded; if (say) a financial services company can determine that the details for a particular credit card or bank account are being offered for sale, this provides the opportunity for the account to be 'locked' or deactivated.

It can also be extremely difficult to configure monitoring software to search the Dark Web. Whilst it is technically relatively straightforward to configure systems to be Tor-enabled (although connections are typically rather slow), there are generally no robust indexes of Dark Web content (such as the search engines and zone files used to search surface-web content), not least because the .onion addresses for any given website - which usually consist of long, random alphanumeric strings - are generally short-lived and change over time. A number of Dark Web search engines do exist, together with ad-hoc indexes of Dark Web content posted by users on sites such as Pastebin, but the information on these sources typically becomes out-of-date rather quickly.

The nature of the content on the Dark Web also means that security concerns can be an issue for brand-protection service providers wishing to build their capabilities in this area.

7. Mobile-based technologies

As Internet engagement has continued to grow over recent years, an increasing proportion of Internet use is conducted over mobile devices[14,15], using a wide ecosystem of mobile apps. Many platforms are now almost exclusively mobile-based, often with little or no corresponding web presence - popular examples might include the WeChat / Weixin platforms, public groups on messaging services such as WhatsApp, and e-commerce platforms such as Pinduoduo. Many brand-protection service providers use legacy monitoring technologies which were designed specifically for analysing HTML content on the regular Internet and are often poorly equipped to address mobile technologies. In some cases, the work-around is to make use of standalone mobile devices or emulators - on which significant proportions of the monitoring is conducted manually - and there typically remains significant work to be done in order to fully integrate the relevant technologies into core monitoring capabilities.

8. Addressing the Web3 landscape

Web3 (also known as 'Web 3.0') is a general term referring to decentralised content on the Internet, with a particular focus on blockchain technologies. Blockchains are publicly accessible digital ledgers in which transactions are recorded, and form the basis of many digital currencies (or 'cryptocurrencies') (such as Bitcoin), in addition to a number of other applications, such as supply-chain control by brand owners. From a brand-protection viewpoint, the main related areas of interest are typically NFTs and blockchain domains[16].

NFTs (non-fungible tokens) are digital files whose ownership is recorded on a blockchain. They are most commonly associated with graphics files (such as artworks and branded imagery) or other types of digital content (such as audio or music files). However, brand owners are increasingly incorporating NFTs into their business models, including areas such as the production and trade of virtual branded items (e.g. items to be worn by avatars in virtual-reality environments within the 'metaverse', the name given to a generalised connected environment of 3D virtual worlds). Consequently, unofficial branded NFTs can be a source of concern for brand owners.

Blockchain domains - which are recorded (together with their ownership details) on a blockchain, rather than using traditional registrars and web hosting - have a number of similarities to 'classic' domain names, and can be utilised in a number of ways. The most common uses are the creation of decentralised websites on peer-to-peer (P2P) platforms, to be accessed via specially-enabled browsers, or as addresses for sending and receiving cryptocurrency. However, the blockchain domain ecosystem is essentially unregulated, and nothing analogous to domain-name zone files is available. The system is made additionally more complicated by the fact the infrastructure allows for the possibility of domain-name 'clashes' - i.e. the potential for the same name to exist independently on distinct blockchains. As with traditional domain names, blockchain domains with brand-specific names can be threat to brand owners, and a potential source of confusion for customers.

Both NFTs and blockchain domains can be traded on NFT marketplaces (such as OpenSea), and the monitoring of these sites is typically the primary source of intelligence utilised by those brand-protection service providers offering capabilities in this area. For blockchain domains particularly, this approach is less than satisfactory, and offers nothing approaching the sort of comprehensive coverage as is available for regular gTLD domain names via zone-file analysis. Some additional information on the existence of registered blockchain domains is typically available through direct searches within databases provided by blockchain domain registrars and nameserver providers; however, the problem of more comprehensive detection is much more difficult to solve, potentially involving analysis of the content of the individual blockchains directly.

Another difficulty to be overcome in service offerings relating to NFTs and blockchain domains is the issue of enforcement against infringing content. In some cases, enforcement can be carried out through the submission of a DMCA (Digital Millennium Copyright Act) notice, and some NFT marketplaces have specific takedown procedures for content which infringes protected IP. However, in many cases, this simply involves the item being 'delisted' from the marketplace in question. In the future, we may see a move towards more rigorous enforcement, potentially involving forced transfers of ownership. Part of the problem is that the legal issues surrounding NFTs and blockchain domains are, in many cases, still not well-defined and are rapidly evolving, complicated by factors such as the fact that ownership of an NFT ownership does not necessarily grant ownership of copyright for the embedded content.

Beyond #8: Other emerging technologies

As new Internet technologies continue to emerge and develop, they will bring with them new risks for brand owners and associated challenges for brand-protection service providers, who will need to continue to observe and innovate in order to stay ahead of the curve.

At any given time, it is unclear where the next area of concern will come from. Currently, there is a great deal of buzz and speculation about artificial intelligence (AI) technologies and chatbots such as ChatGPT, but it is less obvious how these may affect brand-protection considerations. In this context, I am referring to content associated with, or produced by, AI applications. (Conversely, however, it seems highly likely that AI capabilities will be increasingly built into technologies used to facilitate the brand-protection process - i.e. tools to assist with monitoring, prioritisation, clustering and enforcement.)

Users are able to communicate with AI technologies such as ChatGPT via natural language, which are then able to construct responses based on information with which they have been 'trained'. This means that the information available from a chatbot is only as good as the data with which it has been trained (essentially, in the case of ChatGPT, including large volumes of Internet databases[17,18]), and should really be treated with at least as much caution as the old "I'm Feeling Lucky" button on Google, where the user is just presented with a single response (not necessarily the most reliable one!) to any given query. This point is all the more valid given the ability of chatbots to extrapolate, and provide responses based on incomplete information. What this all means is that chatbots pose the risk of providing information about (say) a company or brand which is misleading or otherwise damaging to corporate reputation. However, since responses are generated dynamically in response to queries (rather than being 'fixed', as in the content of an HTML webpage), it is not clear how these issues might be addressed from a brand-protection point of view. Further complications surround issues such as the ownership of rights to content produced by AI technologies[19].

Where chatbots may be of particular concern from a brand-protection and cybersecurity point of view is in their ability to rapidly create content of a wide variety of types, in a range of different styles - including the ability to write and de-bug computer code. What this may mean is that the entry barrier for infringers wishing to create compelling phishing e-mails[20], or write malicious programs ('malware')[21] may be significantly diminished. The likelihood is - at least in the first generations of AI technologies - that AI will not so much change the types of attack which are possible, but rather the ease with which they can be executed[22].

Another issue surrounds use-cases in which AI systems are 'trained' with confidential corporate information as part of the process of creation of company materials (such as marketing releases). These scenarios raise the possibility for the information to be accessed by third parties, either directly via hacking, or via content included in the responses provided to other users, depending on the ways in which information is 'shared' within the infrastructure of the AI technology itself[23]

References

[1] https://www.cst.cam.ac.uk/ring/halloffame

[2] https://www.markmonitor.com/download/ds/MarkMonitor-Corporate-Overview.pdf

[3] https://www.claymath.org/millennium-problems

[4] https://www.linkedin.com/pulse/assessing-mediating-digital-risk-landscape-brand-david-barnett/

[5] https://www.worldtrademarkreview.com/global-guide/anti-counterfeiting-and-online-brand-enforcement/2022/article/creating-cost-effective-domain-name-watching-programme

[6] https://www.cscdbs.com/blog/branded-domains-are-the-focal-point-of-many-phishing-attacks/

[7] https://www.cscdbs.com/en/resources-news/threatening-domains-targeting-top-brands/

[8] https://www.linkedin.com/pulse/hyphenated-domain-infringements-david-barnett/

[9] https://newgtlds.icann.org/en/about/program

[10] https://www.cscdbs.com/blog/the-world-of-the-subdomain/

[11] https://www.linkedin.com/pulse/exploring-domain-hostname-based-infringements-david-barnett/

[12] https://www.cscdbs.com/blog/do-you-see-what-i-see-geotargeting-in-brand-infringements/

[13] 'Brand Protection in the Online World: A Comprehensive Guide' by David Barnett (2016). Chapter 11: ''Deep' and 'Dark' Web'

[14] https://www.statista.com/statistics/617136/digital-population-worldwide/

[15] https://www.linkedin.com/pulse/holistic-brand-fraud-cyber-protection-using-domain-threat-barnett/

[16] https://www.linkedin.com/pulse/rise-nft-david-barnett

[17] https://www.sciencefocus.com/future-technology/gpt-3/

[18] https://techcrunch.com/2023/03/23/openai-connects-chatgpt-to-the-internet/

[19] https://intellectual-property-helpdesk.ec.europa.eu/news-events/news/intellectual-property-chatgpt-2023-02-20_en

[20] https://securityboulevard.com/2023/01/what-does-chat-gpt-imply-for-brand-impersonation-qa-with-dr-salvatore-stolfo/

[21] https://www.digitaltrends.com/computing/chatgpt-created-malware/

[22] https://venturebeat.com/security/security-risks-evolve-with-release-of-gpt-4/

[23] https://blogs.blackberry.com/en/2023/04/is-chatgpt-safe-for-organizations-to-use

This article was first published on 25 May 2023 at:

https://circleid.com/posts/20230525-the-millennium-problems-in-brand-protection

Friday, 3 March 2023

Developing a methodology for benchmarking marketplace brand infringements

Introduction

One of the primary aims of a brand-protection programme is typically the ability to determine the extent of brand infringements on e-commerce marketplaces - and ideally, to be able to benchmark this metric against comparable competitor brands. In this article, I discuss a simple initial possible methodology for quantifying this characteristic, based on the price point of the items in the listings returned in response to a brand-specific search (with a low price point typically indicating that a listing may be of interest). 

The methodology considers the first page of results returned on any given marketplace, in response to a relevant search, and attempts to quantify the proportion of infringing listings within this dataset - a concept which is familiar from other areas in which metrics for measuring infringements or brand-protection effectiveness are required (e.g. where one aim of a brand-protection programme might be to 'clean up' the first page of results, so that only legitimate products or sellers are returned, and no infringing products are present). 

On marketplaces, listings of potential interest can typically fall into a range of categories, including counterfeit goods, trademark infringements, compatible items, 'grey-market' trade (i.e. legitimate goods sold outside approved channels), legitimate second-hand goods, and so on. Attempting to quantify the overall level of infringements based purely on price point will always therefore have shortcomings, and it may also be necessary to apply some degree of 'filtering' in order to obtain meaningful results and compare like with like. Of course, in practice, any definitive determination of infringement type will always require more detailed manual analysis for each listing (potentially also combined with other factors such as test purchases). However, in this article, I consider a high-level approach which may be at least partially automatable.

Since a simple search for just a brand name will be likely to return a mix of product types (with an associated range of prices, even for the legitimate items), I take the approach of considering one or more specific products for each brand, each of which will have a single, well-defined price for the legitimate item.

Exploring a test case

In this investigation, I consider the iPhone 14 Pro - an example of a relatively new, high-desirability product of a type which is typically prone to counterfeits and other infringement issues. I consider the listings returned on the first page of results of a specific marketplace on 01-Mar-2023, approximately six months after the initial official release of the product[1]

Since one of the aims of the analysis is to be able to benchmark against other brands and products, it will be beneficial to compare the marketplace listing prices against the actual price of the genuine product, rather than just considering the absolute numbers. This can be achieved by expressing the price per item in the marketplace listing as a proportion of the genuine item list price ($999 in the case of the iPhone 14 Pro[2]) - a measure I refer to as 'relative price'.

Where appropriate, it may also be necessary to apply a product-type filter to the results provided by the marketplace - for example, a simple search for 'iPhone 14 Pro' may return a mix of product types (including phones, accessories, and so on); however, setting the product type filter to 'mobile phones' specifically will (in theory) only return listings for phones themselves, so the spread of prices across the listings will be more reflective of the types of infringements present across the dataset of mobile-phone results. 

Even then, a low price point (say) is, in itself, not necessarily indicative that a listing represents a counterfeit product. Other types of listings (such as legitimate second hand-products and trademark infringements - e.g. where a brand name is used in the listing title so as to attract search traffic to the listing, but the listing itself is for a third-party branded product) may also be associated with low prices. The first possibility can be mediated to some degree by considering only specific marketplaces where the extent of second-hand trade is limited (e.g. B2B marketplaces)[3]; conversely, separating out (say) counterfeits from other infringement types is much more difficult using only a largely-automated, price-based approach. However, the argument can be made that all such listings (with low price points) are likely to be infringing in some way - all we are therefore looking to do is quantify the overall size of this general infringement landscape.

As an illustration of the results, shown below (Table 1) is an overview of the top ten results returned in response to a listing for 'iPhone 14 Pro' on the marketplace in question, with the product filter set to 'mobile phones'.

Title

                                                                
Price per item
(min. listed) ($)

                          
Min. order quantity Quantity
(max. listed)

                     
Brand name
                         
Relative price
  Wholesale mobile phone Original Smart
  5G Mobile Cell Phones for iphone 11
  128GB
50.00 2 300   for Apple 0.050
  New Arrival Original Brand Phone 11pro
  max 12 mini Waterproof Face
  Recognition 256gb 512gb 1TB Game
  Mobile Phone for iPhone 13
399.00 2 50   original 0.399
  Smartphone mobile iphone 11ProMax
  256gb 5g usa spec original no scratches
  body low price for wholesale 6.5inch
  screen game phone
497.00 1 1   Other 0.497
  Hot Selling PHONE 14 PRO MAX 12GB+
  512GB 6.7 Inch full Display Android
  10.0 Mobile Phone I13 PRO MAX Cell
  Phone Smartphone
27.72 1 99,999   Android
  Smartphone
0.028
  Low price wholesale smartphone 14
  Pro Max 8GB+256GB 7.3in 8core 4G
  LET global Edition smartphone
72.00 1 1,000   W&O 0.072
  New Global I 14 Pro Max Cell Phone 7.3
  Inch Big Screen 5G Smartphone 16GB +
  1TB Global Unlock Dual SIM Android
  Mobile Phone
95.00 1 10,000   Other 0.095
  Free shipping phone I13 pro max 8GB+
  256GB 6.7 Inch full Display Android 10.0
  Mobile Phone PHONE13 PRO MAX Cell
  Phone Smartphone
66.00 1 99,999   Smartphone
  S22
0.066
  i13 Pro cash on delivery mobile phone 8+
  16MP New Original Unlocked Smartphone
  6.8" Display OEM
66.00 1 99,999   Android
  Smartphone
0.066
  High Quality i 14 Pro Max 5G 6.8 Inch
  Original Mobile Phone 16GB+1TB Large
  Memory Smart Phone Beauty Camera
  Gaming Cellphone
32.00 1 20,000   Other 0.032
  High Quality i 14 Pro Max 5G 6.8 Inch
  Original Mobile Phone 16GB+1TB Large
  Memory Smart Phone Beauty Camera
  Gaming Cellphone
55.10 1 20,000   Other 0.055

Table 1: Details of top ten listings returned in response to a search for 'iPhone 14 Pro', with the product filter set to 'mobile phones'

Notes:

  • The 'price per item' is given as the lowest price referenced in the listing, in cases where the unit price may vary dependent on the quantity offered.
  • The 'quantity (max. listed)' value is given as either the maximum quantity stated as being available, or the maximum quantity for which a unit price is specified in the listing (whichever is greater).

From the total set of (48) listings returned on the first page of results, a number of other observations are particularly noteworthy:

  1. None of the listings contains what would normally be described as a counterfeit product; none shows Apple branding in the product image, and none cites the product brand name as Apple (with the exception of the first listing, in which the product is stated as 'for Apple'; this is a technique commonly used by sellers to describe compatible products - although this would be largely non-sensical for a smartphone listing - or as a means of circumventing enforcement efforts), though some listings do give the brand as 'original' or 'OEM'. However, many of the listings in the dataset would constitute trademark infringements, with brand names given in several cases as 'other', 'Android smartphone', or a brand name referring specifically to the seller in question.

  2. Several of the returned listings appear not to be infringing the iPhone 14 Pro product in any way, as the marketplace seems to also return a number of listings referring only to one or more of the individual keywords in the search phrase. Only a subset of the results (those referring explicitly to '14pro' or 'phone14' (both with or without spaces) - shown in bold text in Table 1) are likely to be directly infringing. Accordingly, when carrying out the price-point analysis, it will be beneficial to apply some filtering in order to exclude all except these listings. 

  3. It is also informative that a number of the relevant listings do make reference to 'i 14 pro' or 'I14 pro' - presumably as a way of avoiding directly infringing the iPhone brand name, and potentially also aiming to circumvent detection. Use of brand variations of this type is popular with infringers.

  4. Amongst the listings, a range of maximum quantities (per listing) was observed, from 1 to 10,000,000.

  5. All except one of the listings are for sellers based in China, with a significant number operating out of the manufacturing centres of Shenzhen, a trend which has frequently been observed for sellers of infringing products.

For the 48 listings, the distribution of relative price per item is as shown in Figure 1.

Figure 1: Distribution of relative price per item, for the full set of 48 listings returned on the first page of results in response to a search for 'iPhone 14 Pro', with the product filter set to 'mobile phones'

The results show strikingly that the listings are dominated by products at a very low price point, with the vast majority of items at 10% of list price or lower (i.e. ≤ 0.10 relative price). 

It is instructive to consider some examples of the listings with the lowest price point (excluding the non-'14pro' and non-'phone14' results, as discussed in point (2) above), to analyse the types of infringement present. Of the five listings with the lowest relative prices, for example, all are offering high quantities of items (up to between 10,000 and 99,999), and all appear to represent trademark infringements (or potentially to be involved in the supply chain for counterfeit products) (Figure 2).

Figure 2: Examples of listings with very low price points, both offering 'customized logo' and 'customized packaging' for bulk orders

For developing this methodology further, it is advantageous to express the number of listings in each relative price 'bin' as a proportion of the total, rather than as an absolute number. This has a couple of advantages, specifically:

  • It allows for easier comparison across different marketplaces, where the number of results returned by page may differ.
  • It allows filtering of results to remove any 'false positives' (as discussed in point (2) above).

It also simplifies the calculations if the bins are a consistent width throughout (in this case, 0.02).

This therefore gives the results for the iPhone 14 Pro search for the marketplace in question in the format shown in Figure 3 below (in which the 20 non-relevant listings have been excluded).

Figure 3: Distribution of relative price per item, for the set of listings returned on the first page of results in response to a search for 'iPhone 14 Pro', with the product filter set to 'mobile phones', and with non-relevant / non-infringing listings excluded

In order to carry out the benchmarking across different brands or marketplaces, it is also useful to construct a single metric (or value) which provides a measure of the distribution of relative price points across the set of listings. Essentially, we would like this number to represent the proportion of the 'area under the graph' at the low-price-point end of the relative price distribution chart. 

This can be achieved by summing up the heights of the individual columns, but weighting more heavily (i.e. applying a larger multiplying factor(s) to the heights of) the columns at the low-price end. The weightings can be selected in a number of different ways; one possible methodology is to calculate a weighting which is inversely proportional to the relative price value at the mid-point of the bin in question (such that, for example, the height of the column for the bin associated with a price-point range of 0.00 to 0.02 - i.e. with a mid-point of 0.01 - is weighted by a factor of (1/0.01), or 100).

This methodology allows us to calculate a single price-point metric (P)[4], whose value increases according to the proportion of the listings in the sample associated with lower price points (and therefore provides a measure of the potential scale of the infringement landscape). In the case where all listings have a relative price of 1.00 (i.e. potentially just a set of legitimate product listings), the value of P will be 1[5]. In this case, for the distribution of iPhone 14 Pro listing price-points shown in Figure 3, the value of P is 19.180.

The approach thereby allows us to benchmark the product against other products, brands, or marketplaces. For example, considering the same marketplace, but looking instead at the comparable Galaxy S23 Ultra product (RRP = $1199.99)[6], and similarly applying filtering to remove non-relevant listings, we obtain the price distribution as shown in Figure 4.

Figure 4: Distribution of relative price per item, for the set of listings returned on the first page of results in response to a search for 'Galaxy S23 Ultra', with the product filter set to 'smart phones', and with non-relevant / non-infringing listings excluded

In this case, the price-point metric value (P) is 21.643, indicating a greater proportion of listings at the lowest price points than for the iPhone product on the same marketplace, and potentially therefore a larger infringement landscape. This is consistent with what we can subjectively see in Figure 4, with a greater peak in the lowest occupied relative-price bin (between values of 0.02 and 0.04).

Conclusion

Whilst taking a very simple-minded approach, the methodology discussed above does provide a basic measure of the proportion of listings in a set of marketplace results which show low price points and, by extension, a measure of the potential scale of the infringement landscape. Obviously this assertion is only valid if we accept low price point as a proxy for a listing to be of interest, but previous analysis has certainly shown that it is at least one such valid indicator (and as also borne out by the examples presented in this article).

In practice, calculation of the price-point metric could be automatable, based on collection of marketplace data using monitoring tools, combined with (a) scraping technology to automatically extract the price information and (b) filtering technology to remove false positives in the results.

By applying and expanding these ideas, it would be possible to carry out cross-brand and cross-marketplace benchmarking, and potentially to track trends in the infringement landscape over time (e.g. in conjunction with an active brand-protection programme of monitoring and enforcement).

Acknowledgements

Thanks must go to Angharad Baber, Irene Oh, Agnes Czolnowska and David Franklin for their feedback and input into this article.

References

[1] https://www.apple.com/newsroom/2022/09/apple-introduces-iphone-14-and-iphone-14-plus/

[2] https://www.apple.com/iphone/

[3] Similarly, data will be less likely to be meaningful on (for example) auction-based marketplaces, particularly in the early stages of auctions when the price point is likely to be low by definition.

[4] Formally, PSi [ (1/mi) × i ], where mi is the relative price at the mid-point of the ith bin, and i is the proportion of listings in that bin.

[5] Approximately(!), depending on where the price-bin boundaries are selected to be.

[6] https://www.samsung.com/us/smartphones/galaxy-s23-ultra/buy/galaxy-s23-ultra-256gb-t-mobile-sm-s918uzgaxau/

This article was first published on 3 March 2023 at:

https://www.linkedin.com/pulse/developing-methodology-benchmarking-marketplace-brand-david-barnett/

A browse around the “.shop”s

This article explores the ecosystem of domains hosted on the .shop ('dot-shop') extension, which is popularly used for e-commerce we...