Showing posts with label counterfeit. Show all posts
Showing posts with label counterfeit. Show all posts

Friday, 3 March 2023

Developing a methodology for benchmarking marketplace brand infringements

Introduction

One of the primary aims of a brand-protection programme is typically the ability to determine the extent of brand infringements on e-commerce marketplaces - and ideally, to be able to benchmark this metric against comparable competitor brands. In this article, I discuss a simple initial possible methodology for quantifying this characteristic, based on the price point of the items in the listings returned in response to a brand-specific search (with a low price point typically indicating that a listing may be of interest). 

The methodology considers the first page of results returned on any given marketplace, in response to a relevant search, and attempts to quantify the proportion of infringing listings within this dataset - a concept which is familiar from other areas in which metrics for measuring infringements or brand-protection effectiveness are required (e.g. where one aim of a brand-protection programme might be to 'clean up' the first page of results, so that only legitimate products or sellers are returned, and no infringing products are present). 

On marketplaces, listings of potential interest can typically fall into a range of categories, including counterfeit goods, trademark infringements, compatible items, 'grey-market' trade (i.e. legitimate goods sold outside approved channels), legitimate second-hand goods, and so on. Attempting to quantify the overall level of infringements based purely on price point will always therefore have shortcomings, and it may also be necessary to apply some degree of 'filtering' in order to obtain meaningful results and compare like with like. Of course, in practice, any definitive determination of infringement type will always require more detailed manual analysis for each listing (potentially also combined with other factors such as test purchases). However, in this article, I consider a high-level approach which may be at least partially automatable.

Since a simple search for just a brand name will be likely to return a mix of product types (with an associated range of prices, even for the legitimate items), I take the approach of considering one or more specific products for each brand, each of which will have a single, well-defined price for the legitimate item.

Exploring a test case

In this investigation, I consider the iPhone 14 Pro - an example of a relatively new, high-desirability product of a type which is typically prone to counterfeits and other infringement issues. I consider the listings returned on the first page of results of a specific marketplace on 01-Mar-2023, approximately six months after the initial official release of the product[1]. 

Since one of the aims of the analysis is to be able to benchmark against other brands and products, it will be beneficial to compare the marketplace listing prices against the actual price of the genuine product, rather than just considering the absolute numbers. This can be achieved by expressing the price per item in the marketplace listing as a proportion of the genuine item list price ($999 in the case of the iPhone 14 Pro[2]) - a measure I refer to as 'relative price'.

Where appropriate, it may also be necessary to apply a product-type filter to the results provided by the marketplace - for example, a simple search for 'iPhone 14 Pro' may return a mix of product types (including phones, accessories, and so on); however, setting the product type filter to 'mobile phones' specifically will (in theory) only return listings for phones themselves, so the spread of prices across the listings will be more reflective of the types of infringements present across the dataset of mobile-phone results. 

Even then, a low price point (say) is, in itself, not necessarily indicative that a listing represents a counterfeit product. Other types of listings (such as legitimate second hand-products and trademark infringements - e.g. where a brand name is used in the listing title so as to attract search traffic to the listing, but the listing itself is for a third-party branded product) may also be associated with low prices. The first possibility can be mediated to some degree by considering only specific marketplaces where the extent of second-hand trade is limited (e.g. B2B marketplaces)[3]; conversely, separating out (say) counterfeits from other infringement types is much more difficult using only a largely-automated, price-based approach. However, the argument can be made that all such listings (with low price points) are likely to be infringing in some way - all we are therefore looking to do is quantify the overall size of this general infringement landscape.

As an illustration of the results, shown below (Table 1) is an overview of the top ten results returned in response to a listing for 'iPhone 14 Pro' on the marketplace in question, with the product filter set to 'mobile phones'.

Title

                                                                
Price per item
(min. listed) ($)

                          
Min. order quantity Quantity
(max. listed)

                     
Brand name
                         
Relative price
  Wholesale mobile phone Original Smart
  5G Mobile Cell Phones for iphone 11
  128GB
50.00 2 300   for Apple 0.050
  New Arrival Original Brand Phone 11pro
  max 12 mini Waterproof Face
  Recognition 256gb 512gb 1TB Game
  Mobile Phone for iPhone 13
399.00 2 50   original 0.399
  Smartphone mobile iphone 11ProMax
  256gb 5g usa spec original no scratches
  body low price for wholesale 6.5inch
  screen game phone
497.00 1 1   Other 0.497
  Hot Selling PHONE 14 PRO MAX 12GB+
  512GB 6.7 Inch full Display Android
  10.0 Mobile Phone I13 PRO MAX Cell
  Phone Smartphone
27.72 1 99,999   Android
  Smartphone
0.028
  Low price wholesale smartphone 14
  Pro Max 8GB+256GB 7.3in 8core 4G
  LET global Edition smartphone
72.00 1 1,000   W&O 0.072
  New Global I 14 Pro Max Cell Phone 7.3
  Inch Big Screen 5G Smartphone 16GB +
  1TB Global Unlock Dual SIM Android
  Mobile Phone
95.00 1 10,000   Other 0.095
  Free shipping phone I13 pro max 8GB+
  256GB 6.7 Inch full Display Android 10.0
  Mobile Phone PHONE13 PRO MAX Cell
  Phone Smartphone
66.00 1 99,999   Smartphone
  S22
0.066
  i13 Pro cash on delivery mobile phone 8+
  16MP New Original Unlocked Smartphone
  6.8" Display OEM
66.00 1 99,999   Android
  Smartphone
0.066
  High Quality i 14 Pro Max 5G 6.8 Inch
  Original Mobile Phone 16GB+1TB Large
  Memory Smart Phone Beauty Camera
  Gaming Cellphone
32.00 1 20,000   Other 0.032
  High Quality i 14 Pro Max 5G 6.8 Inch
  Original Mobile Phone 16GB+1TB Large
  Memory Smart Phone Beauty Camera
  Gaming Cellphone
55.10 1 20,000   Other 0.055

Table 1: Details of top ten listings returned in response to a search for 'iPhone 14 Pro', with the product filter set to 'mobile phones'

Notes:

  • The 'price per item' is given as the lowest price referenced in the listing, in cases where the unit price may vary dependent on the quantity offered.
  • The 'quantity (max. listed)' value is given as either the maximum quantity stated as being available, or the maximum quantity for which a unit price is specified in the listing (whichever is greater).

From the total set of (48) listings returned on the first page of results, a number of other observations are particularly noteworthy:

  1. None of the listings contains what would normally be described as a counterfeit product; none shows Apple branding in the product image, and none cites the product brand name as Apple (with the exception of the first listing, in which the product is stated as 'for Apple'; this is a technique commonly used by sellers to describe compatible products - although this would be largely non-sensical for a smartphone listing - or as a means of circumventing enforcement efforts), though some listings do give the brand as 'original' or 'OEM'. However, many of the listings in the dataset would constitute trademark infringements, with brand names given in several cases as 'other', 'Android smartphone', or a brand name referring specifically to the seller in question.

  2. Several of the returned listings appear not to be infringing the iPhone 14 Pro product in any way, as the marketplace seems to also return a number of listings referring only to one or more of the individual keywords in the search phrase. Only a subset of the results (those referring explicitly to '14pro' or 'phone14' (both with or without spaces) - shown in bold text in Table 1) are likely to be directly infringing. Accordingly, when carrying out the price-point analysis, it will be beneficial to apply some filtering in order to exclude all except these listings. 

  3. It is also informative that a number of the relevant listings do make reference to 'i 14 pro' or 'I14 pro' - presumably as a way of avoiding directly infringing the iPhone brand name, and potentially also aiming to circumvent detection. Use of brand variations of this type is popular with infringers.

  4. Amongst the listings, a range of maximum quantities (per listing) was observed, from 1 to 10,000,000.

  5. All except one of the listings are for sellers based in China, with a significant number operating out of the manufacturing centres of Shenzhen, a trend which has frequently been observed for sellers of infringing products.

For the 48 listings, the distribution of relative price per item is as shown in Figure 1.

Figure 1: Distribution of relative price per item, for the full set of 48 listings returned on the first page of results in response to a search for 'iPhone 14 Pro', with the product filter set to 'mobile phones'

The results show strikingly that the listings are dominated by products at a very low price point, with the vast majority of items at 10% of list price or lower (i.e. ≤ 0.10 relative price). 

It is instructive to consider some examples of the listings with the lowest price point (excluding the non-'14pro' and non-'phone14' results, as discussed in point (2) above), to analyse the types of infringement present. Of the five listings with the lowest relative prices, for example, all are offering high quantities of items (up to between 10,000 and 99,999), and all appear to represent trademark infringements (or potentially to be involved in the supply chain for counterfeit products) (Figure 2).

Figure 2: Examples of listings with very low price points, both offering 'customized logo' and 'customized packaging' for bulk orders

For developing this methodology further, it is advantageous to express the number of listings in each relative price 'bin' as a proportion of the total, rather than as an absolute number. This has a couple of advantages, specifically:

  • It allows for easier comparison across different marketplaces, where the number of results returned by page may differ.
  • It allows filtering of results to remove any 'false positives' (as discussed in point (2) above).

It also simplifies the calculations if the bins are a consistent width throughout (in this case, 0.02).

This therefore gives the results for the iPhone 14 Pro search for the marketplace in question in the format shown in Figure 3 below (in which the 20 non-relevant listings have been excluded).

Figure 3: Distribution of relative price per item, for the set of listings returned on the first page of results in response to a search for 'iPhone 14 Pro', with the product filter set to 'mobile phones', and with non-relevant / non-infringing listings excluded

In order to carry out the benchmarking across different brands or marketplaces, it is also useful to construct a single metric (or value) which provides a measure of the distribution of relative price points across the set of listings. Essentially, we would like this number to represent the proportion of the 'area under the graph' at the low-price-point end of the relative price distribution chart. 

This can be achieved by summing up the heights of the individual columns, but weighting more heavily (i.e. applying a larger multiplying factor(s) to the heights of) the columns at the low-price end. The weightings can be selected in a number of different ways; one possible methodology is to calculate a weighting which is inversely proportional to the relative price value at the mid-point of the bin in question (such that, for example, the height of the column for the bin associated with a price-point range of 0.00 to 0.02 - i.e. with a mid-point of 0.01 - is weighted by a factor of (1/0.01), or 100).

This methodology allows us to calculate a single price-point metric (P)[4], whose value increases according to the proportion of the listings in the sample associated with lower price points (and therefore provides a measure of the potential scale of the infringement landscape). In the case where all listings have a relative price of 1.00 (i.e. potentially just a set of legitimate product listings), the value of P will be 1[5]. In this case, for the distribution of iPhone 14 Pro listing price-points shown in Figure 3, the value of P is 19.180.

The approach thereby allows us to benchmark the product against other products, brands, or marketplaces. For example, considering the same marketplace, but looking instead at the comparable Galaxy S23 Ultra product (RRP = $1199.99)[6], and similarly applying filtering to remove non-relevant listings, we obtain the price distribution as shown in Figure 4.

Figure 4: Distribution of relative price per item, for the set of listings returned on the first page of results in response to a search for 'Galaxy S23 Ultra', with the product filter set to 'smart phones', and with non-relevant / non-infringing listings excluded

In this case, the price-point metric value (P) is 21.643, indicating a greater proportion of listings at the lowest price points than for the iPhone product on the same marketplace, and potentially therefore a larger infringement landscape. This is consistent with what we can subjectively see in Figure 4, with a greater peak in the lowest occupied relative-price bin (between values of 0.02 and 0.04).

Conclusion

Whilst taking a very simple-minded approach, the methodology discussed above does provide a basic measure of the proportion of listings in a set of marketplace results which show low price points and, by extension, a measure of the potential scale of the infringement landscape. Obviously this assertion is only valid if we accept low price point as a proxy for a listing to be of interest, but previous analysis has certainly shown that it is at least one such valid indicator (and as also borne out by the examples presented in this article).

In practice, calculation of the price-point metric could be automatable, based on collection of marketplace data using monitoring tools, combined with (a) scraping technology to automatically extract the price information and (b) filtering technology to remove false positives in the results.

By applying and expanding these ideas, it would be possible to carry out cross-brand and cross-marketplace benchmarking, and potentially to track trends in the infringement landscape over time (e.g. in conjunction with an active brand-protection programme of monitoring and enforcement).

Acknowledgements

Thanks must go to Angharad Baber, Irene Oh, Agnes Czolnowska and David Franklin for their feedback and input into this article.

References

[1] https://www.apple.com/newsroom/2022/09/apple-introduces-iphone-14-and-iphone-14-plus/

[2] https://www.apple.com/iphone/

[3] Similarly, data will be less likely to be meaningful on (for example) auction-based marketplaces, particularly in the early stages of auctions when the price point is likely to be low by definition.

[4] Formally, P = Si [ (1/mi) × ℓi ], where mi is the relative price at the mid-point of the ith bin, and ℓi is the proportion of listings in that bin.

[5] Approximately(!), depending on where the price-bin boundaries are selected to be.

[6] https://www.samsung.com/us/smartphones/galaxy-s23-ultra/buy/galaxy-s23-ultra-256gb-t-mobile-sm-s918uzgaxau/

This article was first published on 3 March 2023 at:

https://www.linkedin.com/pulse/developing-methodology-benchmarking-marketplace-brand-david-barnett/

Thursday, 9 February 2023

Calculation of return on investment for brand-protection programmes: Thoughts towards a new paradigm

Pre-existing ideas

Numerous previous studies have considered methodologies for calculating the return on investment (ROI) of brand-protection programmes which incorporate components of monitoring and enforcement. These ideas can be important both to justify the spend on a programme in the first place, and to assess its impact once established. Correspondingly, 'classic' ROI calculations can be categorised into two main types: the first (known as 'a priori' calculations) consider the probable infringement landscape in advance of the implementation of a brand-protection programme; the second aims to quantify the actions taken as part of an active enforcement initiative[1]. It is the latter category with which we are primarily concerned in this article.

To a very high level, many ROI calculation methodologies use a formulation along the lines of:

R = C × E

where R, the ROI (within a given timeframe) (i.e. the benefit of the brand-protection programme, to be offset against the associated spend) is equal to the product of C, the 'cost' of a pre-existing infringement being active, and E, the number of infringements removed through enforcement as part of the brand-protection programme (in the same timeframe).

Very many assumptions are typically required in order to estimate these figures. In some methodologies, the assumed 'cost' associated with a live infringement may be reflective of an estimate of its direct financial impact (e.g. the typical loss from a phishing incident); in others it may be calculated as the proportion of lost revenue which is reclaimable following deactivation of the infringement (i.e. the 'cost' in the above formulation essentially reflecting the pre-enforcement impact of not yet having taken the infringement down). In these types of approaches, it is very rare that these figures can be measured directly and therefore a number of assumptions (or 'proxies' for the data) are required. In cases of domain acquisition, for example, it may be appropriate to make use of figures such as web traffic when quantifying impact; for marketplace listings, it is typically necessary to consider factors such as price and quantity of items in the listings removed. In both cases, the methodology needs to consider assumed conversion rates (i.e. the proportion of customers who can be 'monetised' by the brand owner - e.g. those who will make a legitimate purchase once the source of infringements is removed)[2,3]. Even this part of the process is far from simple; complications include factors such as: 

  1. The conversion rate will be (strongly) dependent on the nature and price of the item (e.g. it will be much lower for (say) an obvious counterfeit, such as an item passing off as a high-end luxury brand but with a very low price point)[4].

  2. The conversion rate for customers knowingly navigating to an official brand website will potentially be different to that for those Internet users intending to visit a third-party standalone e-commerce site (if we are considering the case where this domain may subsequently have been acquired by the brand owner and its traffic re-directed to their official site) - this consideration involves taking account of a principle sometimes referred to as the 'substitution effect'[5]. 

Alternative proxies for the above figures may also need to be utilised, depending on the web channel under consideration (e.g. where absolute estimates of web traffic are not available or appropriate). For example, on social media, the 'exposure' or 'reach' of content can be estimated using numbers of 'likes' or followers; for mobile apps, the number of downloads may be relevant; for file sharing, it may be appropriate to consider the number of individuals accessing the content (e.g. 'seeds' and 'leechers' for BitTorrent). 

Numerous other approaches can also be taken. The ultimate objective when estimating the 'value' of a website is the identification a direct measure of the revenue it generates (e.g. via direct sales of products, for an e-commerce site). In practice, this information is almost never publicly available, though it is sometimes possible to make estimations via shipping or logistics information available through third-party databases. Some methodologies will utilise web-analytics tools to estimate value based on factors such as advertising spend by the site owner, or will analyse outgoing site traffic (e.g. to payment service provider platforms) to estimate customer volume and/or conversion rates[6]. 

It has also previously been noted that sometimes determination of ROI can reflect more qualitative goals (i.e. the statements of 'what success looks like' for a brand-protection programme). For example, a brand owner may consider a programme 'successful' once there are no infringing results returned on the first page of search-engine results, or in pages of search results on a range of key marketplace sites, in response to brand-specific queries. Similarly, the 'ownership of the buy button' (i.e. being the first vendor listed for a particular product on an e-commerce marketplace site) might be a key aim.

The success of a brand-protection initiative can also be judged based on other (again, more quantitative) metrics which may be only available to the brand owner themselves (as opposed to, say, a brand-protection service provider partner). These might include factors such as increases in the numbers of visitors to physical stores, or in volumes of traffic to official websites (as might be directly measurable using the brand owner's webserver log information).

Beyond this, wholly different methodologies can also be applied. Some will take account of 'intangible' factors such as brand value[7], considering the spend on brand protection to be a business cost necessary to lower the risk of damage to the brand. This type of approach is also not straightforward - higher levels of abuse can be considered an indicator that the brand is a desirable one, which can actually be reflective of greater brand value. Other factors, such as new product launches, can also affect the visibility of the brand and its likelihood of being targeted, all of which can serve to further complicate the landscape. 

However, in this article, we will primarily consider the simpler approaches discussed in previous work, and look at how they can potentially be modified to better account for the overall impact of a brand-protection programme.

Variations over time in the infringement landscape

Part 1: Single-brand analysis

In this section, we consider an extremely simplified model looking at changes in the infringement landscape over time for a brand, considering in the first instance the example of a newly-launched brand. In this case, the growth in the number of infringements over time might look something like that shown in Figure 1.

Figure 1: Mock-up of the changing infringement landscape over time for a newly-launched brand

The above framework is formulated using a timeframe expressed as numbers of months for convenience, though the timescales observed in practice may vary hugely. There is also a deliberate choice to avoid stating any quantitative numbers for the volumes of infringements, as these will also be dependent on any number of different factors - one brand may see tens or hundreds of infringements; other may see many thousands or more. Beyond these points, the construction of the above trend lines is based on the following scenario:

  • Following the launch of the brand (in month 1), there is a ramp-up in the number of monthly numbers of new infringements ('N') appearing online, up to a constant level.
  • There is also a (slower) ramp-up in the rate of infringements disappearing naturally from the Internet ('natural removal', 'R') even in the absence of any enforcement activity. This will arise through a combination of factors, including: content which is deactivated by the infringer following a period of use; domains expiring after their registration period; older content gradually dropping down search-engine rankings (and potentially therefore eventually ceasing to have any damaging impact), and so on.
  • There is a resulting growth in the cumulative number of active online infringements ('I'), caused by the difference between the monthly values of 'N' and 'R'. 
  • Finally, it seems reasonable to assume in most cases that 'I' will eventually reach a steady state, rather than continuing to grow indefinitely. This implies that 'R' will eventually 'catch up' with 'N' (possibly in part due to the fact that 'N' may also drop off slightly over time, after an initial peak in infringement activity). 

Of course, in practice the exact balance between the above numbers will be dependent on an enormous range of factors, including considerations such as the type of Internet channel. For example, marketplace listings will typically have a shorter 'lifetime' than domain registrations (affecting, for example, the rate at which 'R' catches up with 'N').

Let us now consider the case where a brand-protection programme, incorporating the introduction of enforcement actions for the removal of infringing content, is added into the picture (say, after the landscape has reached steady state in month 12) (Figure 2).

Figure 2: Mock-up of the changing infringement landscape over time, with an enforcement programme introduced in month 12

In this case, we use the following formulation:

  • In month 12, the enforcement programme is introduced, which incorporates a particular level of resource sufficient to action a certain maximum number of takedowns each month. This number will of course need to be greater than the rate at which new infringements appear, if the programme is to be successful. 
  • Following the introduction of the enforcement programme, the rate of natural removal ('R') of infringements will quickly drop off to zero (essentially, the infringements are being removed via enforcement quicker than the rate at which they would otherwise naturally disappear).
  • As enforcement progresses, the cumulative number of infringements drops off from its pre-existing level, until we reach a steady state (the 'whackamole' phase[8]) where the monthly number of enforcements ('E') simply needs to 'keep up' with the rate at which new infringements appear ('N'). In other words, each month a certain number of new infringements appear and these are all removed through the actions of the enforcement programme. (N.B. Equivalently, at this point we could express the 'cumulative number of infringements' ('I') as zero, depending on the point in the month at which we carry out the calculation (i.e. whether pre- or post-enforcement).)

In reality, the situation is likely to be far less straightforward, with a number of additional factors complicating the picture, including (but not limited to) the facts that:

  • The types of infringements actioned over time may change (potentially starting with higher-impact or easier takedowns).
  • Monitoring will inevitably start to uncover lower visibility and/or lower severity infringements once the initial high-visibility, high-impact infringements have been taken down.
  • The rate of appearance of online infringements may change in response to the enforcement programme (e.g. infringers turning their attention to easier targets).
  • The infringers may change their tactics in response to the enforcement programme (e.g. describing goods in different ways) - accordingly, both the monitoring approach and the enforcement methodologies may need to evolve in order to account for this.

Nevertheless, the above very simplistic picture does reflect some of the top-level trends typically seen in a brand-protection programme, with an initial period of 'cleaning up' the pre-existing backlog of infringements followed by a steady-state period of lower required activity, just keeping pace with new infringements as they appear. 

This being the case, we can look to this model to draw insights into how our classic ROI calculation methods could be augmented to provide a fuller picture. In many of the traditional approaches, monthly ROI calculation methodologies make use just of the total monthly numbers of enforcements carried out ('E'). Although the drop-off in the numbers of pre-existing infringements is reflected in the ROI calculations associated with the enforcements carried out during the 'ramp-down' phase itself, it is usually not reflected in the ongoing calculations during the subsequent 'whackamole' phase. Really, it may be preferable to make use of the difference between the ongoing number of infringements ('X') and that observed at the start of the programme ('Y'), if we are to fully assess the impact of the brand-protection programme. In other words, rather than using the number 'X' as the basis of our monthly ROI calculation, it might instead be better to use 'Y – X'. This number instead provides a measure of the value of the ongoing brand-protection programme - essentially, reflecting the difference in the ongoing number of infringements (with the associated 'cost' of them being live) compared with that which would have been observed if the programme were not in place. In practice, determination of these numbers will require the brand-protection initiative to incorporate a comprehensive programme of monitoring (as well as enforcement) throughout, incorporating a full landscape 'audit' at the outset.

Part 2: Benchmarking and the use of controls

To further complicate the situation, what the above approach fails to consider is any changes to the infringement landscape which would have occurred if the brand-protection programme were not being carried out. This is known as the 'attribution' issue in the physical sciences. Of course, once enforcement starts being carried out, we lose the ability to see what would have happened to the numbers of infringements if they were not being actively taken down. It is well established that external factors can significantly change the infringement landscape. For example, numerous previous studies show that real-world events can drive spikes in resulting infringement activity[9]. 

One way in which this problem can be addressed is via comparison with another 'control' brand of a similar type, operating in a similar industry area, but for which brand-protection activity is not being carried out. In practice, a brand owner can never be completely sure what any given competitor is doing, so a more realistic scenario is the use of analysis a group of industry peers, across which the infringement trends over time can be averaged to create a 'benchmark'. Of course, this requires active monitoring across all these brands, and so may be far from straightforward.

In this case, we may end up with a scenario such as that shown in Figure 3, where the control or benchmark brand (actually ideally an average of the data collected across multiple third-party brands) - which we have to assume reflects external drivers in infringement trends in the absence of enforcement initiatives - shows a change in the infringement landscape since the start of the programme for the brand being protected.

Figure 3: Mock-up of the changing infringement landscape over time, with an enforcement programme introduced for the customer brand in month 12, and compared against a (pre-existing, established) benchmark brand(s)

In the above example, the control brand shows a ramp-up in infringements during the period of the brand-protection programme, perhaps driven by an external event of some sort. Additionally, by using a benchmark comprising data from across numerous brands, we reduce the likelihood that the change is driven by some characteristic specific to one brand (such as a new product launch) and increase the likelihood that the change is representative of the industry landscape in general. 

In this case we can assume that, in the absence of a brand-protection programme, the infringement landscape for the customer brand would have increased by the same proportion as that seen for the benchmark brand(s). Therefore, instead of our ROI calculation being a function (' f ') of 'Y – X' (written as 'ROI = f [Y – X]'), we can say that:

ROI = f [ ( (B/A) × Y ) – X ]

Essentially, we are saying that, had the brand-protection programme not been in place, we might have expected the 'background' level of infringements for the customer brand also to have increased by a factor of ('B/A') by the end of the monitoring period, and so the benefit of the programme is in reducing it from this value to the value observed ('X').

Of course, the same approach can also be used if the benchmark shows a decrease in infringements across the monitoring period.

Discussion

The calculation of ROI for brand protection is fiendishly complicated, and no single approach will be applicable in all cases. In any selected methodology, it is necessary to make use of a wide range of assumptions and proxies for the data to which we would ideally like to have access. Nevertheless, there are some general industry-accepted standards for these calculations, many of which utilise metrics around ongoing levels of enforcement activity. In this article, we have considered some approaches which could be taken to modify these methodologies towards a new framework of ideas, involving the following two fundamental changes:

  • Considering the difference between the ongoing levels of enforcement (as a measure of the ongoing level of infringement activity), and those seen at the outset of the programme, as a measure of the overall impact of the brand-protection programme (rather than just considering the ongoing levels of enforcement in their own right).
  • Considering the use of one or (ideally) more benchmark brands, to separate out the observed change in infringement levels (for the customer brand) arising from the enforcement activity, from other background or landscape changes applicable to the industry vertical in general. 

Even then, there are still other factors to consider - the customer brand may also have experienced (company-specific) issues (such as product launches, changes in sales channels or target markets, etc. etc.) which themselves could have driven changes in the number of infringements, even in the absence of an enforcement programme or industry issues. All of this can further complicate the calculations to be carried out.

Additionally, I anticipate that the general philosophy behind ROI calculations may need to evolve further to reflect other issues more directly tied to cybersecurity, as the importance of this area becomes more widely appreciated. A former colleague of mine recently asked in a LinkedIn posting[10]:

""So what's the cost?" is a frequent question I hear. Rather than thinking about the budget required, brands need to consider the financial and reputational costs of repairing the damage when they are impacted."

The key point here is thinking about proactive rather than reactive measures. This issue is particularly relevant when it comes to domain security, where a range of products are available to allow corporations to secure their domains from external attack vectors which can be highly damaging (from both financial and reputational points of view)[11]. The matter is of even greater urgency in a landscape where we still see significant proportions of the world's top companies failing to adequately protect themselves[12]. 

The expected financial loss ('L') per year due to (say) cybersecurity issues (an 'attack') is given[13] by:

L = patt × Catt

where patt is the probability of an attack occurring during the year, and Catt is the financial cost (the 'damage') resulting from the attack. From this, we can say that, if the probability of an attack can be reduced (from pattwithout_security to pattwith_security) through the implementation of domain security measures, the saving ('S') to the organisation can be written as:

S = ( pattwithout_security – pattwith_security ) × Catt

Whilst easy to formulate, this can be much harder to quantify. However, a recent study showed that 88% of organisations were subject to some form of DNS attack in 2021, with each attack costing the enterprise an average of almost $1 million[14]. If, then, the risk of an attack can be (conservatively) reduced from (say) 10% to 1% though the introduction of security measures, this equates to an equivalent annual saving to the company of the order of $90k. If the cost of implementing the security measures is less than this value, the return on investment will be positive. If we factor in also the implications for access to - and cost of - cyberinsurance cover, the importance of domain security products and services becomes ever clearer.

Acknowledgements

Thanks must go to Angharad Baber, Mark Barrett and David Riley for their feedback and input into this article.

References

[1] https://www.worldtrademarkreview.com/anti-counterfeiting/return-investment-proving-protection-pays

[2] https://www.worldtrademarkreview.com/global-guide/anti-counterfeiting-and-online-brand-enforcement/2022/article/creating-cost-effective-domain-name-watching-programme

[3] https://www.cscdbs.com/blog/four-steps-to-an-effective-brand-protection-program/

[4] https://circleid.com/posts/20220726-calculating-the-return-on-investment-of-online-brand-protection-projects

[5] 'Digital Brand Protection: Investigating Brand Piracy and Intellectual Property Abuse' by Steven Ustel (2019). Chapter 17: 'Accounting and Accountability'

[6] 'Digital Brand Protection: Investigating Brand Piracy and Intellectual Property Abuse' by Steven Ustel (2019). Chapter 9: 'Pivots'

[7] https://www.cscdbs.com/blog/brand-abuse-and-ip-infringements/

[8] By 'whackamole' in this context, I am referring to a consistent state in which infringements are reactively taken down as quickly as they appear (rather than implying a random or disordered approach).

[9] https://www.linkedin.com/pulse/four-new-case-studies-domain-registration-activity-spikes-barnett/

[10] https://www.linkedin.com/posts/stuart-fuller-17a7411_what-cisos-can-do-about-brand-impersonation-activity-7027979839747846144-E6ak

[11] https://www.linkedin.com/pulse/holistic-brand-fraud-cyber-protection-using-domain-threat-barnett/

[12] https://www.cscdbs.com/en/resources-news/domain-security-report/ (2022)

[13] This follows from the fact that, mathematically, the expected value ('Ex') of a variable ('X') is given by:

Ex = Si ( p(Xi) × Xi ), where p(Xi) is the probability of X taking the ith value

[14] https://www.efficientip.com/wp-content/uploads/2022/05/IDC-EUR149048522-EfficientIP-infobrief_FINAL.pdf

This article was first published on 9 February 2023 at:

https://www.linkedin.com/pulse/calculation-return-investment-brand-protection-thoughts-david-barnett/

Tuesday, 31 January 2023

Four new case studies of domain registration activity spikes driven by real-world events

Introduction

A variety of previous studies have demonstrated how real-world events can trigger subsequent spikes in domain registrations and infringement activity. Previous CSC articles and reports have focused on issues as diverse as the COVID pandemic[1], the war in Ukraine[2], supply-chain issues affecting the baby-milk and semiconductor industries[3], the Euro 2020 competition[4], the Black Friday and Cyber Monday holiday shopping events[5], and the Reddit stock manipulation campaign targeting the GameStop organisation[6]. 

When a high-impact event or news story takes place, there is typically a resulting burst of public interest and online searches for associated content, and bad actors can take advantage of this 'buzz' for their own gain. There are a number of ways in which this can be implemented, including: the production of content (which can include areas such as the sale of goods via e-commerce sites) relating to the issue at hand; misdirection of users to infringing, unofficial or potentially malicious websites; phishing activity utilising branded domain names to host fraudulent websites or for their e-mail functionality; or monetisation of dormant high-traffic domains through the emplacement of pay-per-click (PPC) links. In some cases, potentially desirable names may also be seized with the intention of subsequent sale to the infringed brand owner (i.e. cybersquatting) or any other interested party. 

In this article, I look at four recent events or news stories, and focus on the manifestation of associated spikes in potential infringements, by considering patterns in domain registration activity. The analysis includes consideration of new registrations ('N'), re-registrations ('R') and domain drops (lapses) ('D').

Findings

Study 1: Changes of UK Prime Minister (Summer 2022)

Summer 2022 was a time of rapid political change in the UK, resulting in two changes of Prime Minister. The associated analysis considers registration activity of domains containing the names of the three leaders, specifically: (i) 'liz' plus 'truss'; (ii) 'rishi' plus 'sunak'; and (iii) 'borisjohnson' (or typos / variations). The findings are shown in Figure 1, where peaks in registration activity can be seen to correspond to associated key news events.

Figure 1: Daily numbers of new registrations ('N') and re-registrations ('R') combined, and dropped ('D') domains, with names relating to the three 2022 UK Prime Ministers (Boris Johnson (top), Liz Truss (middle), Rishi Sunak (bottom)). Key events in the news timeline[7,8] are denoted according to the key shown below.

A: Boris Johnson announces resignation (07-Jul-2022)
B: Liz Truss enters leadership contest (10-Jul-2022)
C: Rishi Sunak frontrunner in leadership contest following second round of voting (13-Jul-2022)
D: Liz Truss confirmed as new Conservative leader and PM following party-member vote (05-Sep-2022)
E: Boris Johnson tenders resignation (06-Sep-2022)
F: Liz Truss faces political rebellion following economic turmoil (04-Oct-2022)
G: Liz Truss announces resignation following appointment of new Chancellor and reversal of 'mini-budget' policies (20-Oct-2022)
H: Rishi Sunak confirmed as new Conservative leader and PM (24-Oct-2022)

In this case, many of the registrations were associated with websites featuring satirical or commentary-related content (Figure 2), though some were of greater concern (misdirection to third-party content or potential phishing activity) (Figure 3). In general, political content can also be of particular concern in cases where it is found to be associated with the spread of misinformation, or be attempting to manipulate voting patterns[9].

Figure 2: Examples of satirical websites identified in the registration dataset - second-level domain names (SLDs) (i.e. the part of the domain name to the left of the dot) are: borisjonson (registered 07-Sep-2022) (top); liztrussgame (registered 23-Oct-2022) (middle); hasrishisunakresignedyet (registered 15-Oct-2022) (bottom)

Figure 3: Examples of other websites identified in the registration dataset - SLDs are: trussliz and wetruzzliz (but displaying content relating to the UK opposition party) (registered 02-Jun-2022) (top); rishisunakforpm (registered 25-Oct-2022) (bottom)

Study 2: FIFA World Cup Qatar 2022

In this study, I consider domain registration activity relating to the 2022 FIFA World Cup competition which took place in Qatar between 20-Nov and 18-Dec 2022. The initial searches focused on all domains containing the keywords 'qatar' or 'world(-)cup', for which over 10,000 registration activity events were identified (comprising 8,690 unique domain names) during a one-year analysis period from December 2021 to December 2022. Continuous activity was identified throughout the year, though unsurprisingly with a ramp-up in new registrations towards the time of the event itself (Figure 4).

Figure 4: Daily (top) and monthly (bottom) numbers of new registrations ('N'), re-registrations ('R'), and dropped ('D') domains with names containing 'qatar' or 'world(-)cup'

In order to take a deeper dive into the highest-relevance domain names, I then focus on searches utilising keywords indicating that the domains under consideration are likely to pertain specifically to the event, rather than just referencing the more generic terms 'Qatar' or 'World Cup'. Specifically, this considers domains with names containing:

  • 'world(-)cup' AND 'qatar'

        OR

  • ['world(-)cup' OR 'qatar'] AND ['football' OR 'futbol' OR 'soccer' OR '2022' OR 'fi(-)fa']

The methodology also considers only those domains which were still active as of the time of analysis (02-Dec-2022) (i.e. those for which the most recent activity event was not a domain drop ('D')). 

This focused analysis yields a dataset of 977 domains, for which the pattern of registration activity (considering only the most recent activity event for each unique domain name) is shown in Figure 5.

Figure 5: Daily (top) and monthly (bottom) numbers of new registrations ('N') and re-registrations ('R') combined, for high-relevance domain names relating to the Qatar World Cup (considering the most recent activity event for each unique domain name)

In this more focused dataset, the overall activity pattern is broadly similar, though an additional peak in registrations is also apparent in early April 2022. This relates to what appears to be one or two specific, short-lived, coordinated registration campaigns of domains with names of the form 'qatar-2022-iX.xyz' and 'worldcup2022-jYYX.buzz' (where 'X' is an additional digit and 'Y' is an additional character). Although none of these domains was found to resolve to any live site at the time of analysis, the .xyz and .buzz new-gTLD domain extensions have been noted as previously being frequently associated with malicious or infringing content[10,11].

Of the 977 high-relevance domains overall, 633 were found to yield an active website response (i.e. an HTTP status code of 200) at the time of analysis. Within this set, a range of (where non-official) potentially infringing or high-threat content types were observed (Figure 6).

Figure 6: Examples of live websites relating to the Qatar World Cup, representing a range of content types of potential concern (with the SLD shown in each case in square brackets) - top to bottom: potential phishing [qatar2022]; piracy [worldcuplivefifa]; gambling [worldcupbet]; ticket sales [qatar-worldcup]; other e-commerce [qatarfootballcup]; cryptocurrency-related [qatarfifaworldcup]; NFT-related [worldcupnft2022]

Study 3: New Year 2023

The new year can be a prime time for brand owners to launch new products, campaigns and marketing activity, and one way in which this can be promoted in a topical fashion is through the registration of new domains making explicit reference to the year. However, similar tactics can also be employed by bad actors, through the registration of desirable domain names. In some cases, these domains may be registered well in advance of the start of the new year itself, as a way of 'getting ahead of the curve'. Accordingly, this study considers activity associated with the registration of domains with names beginning or ending with the string '2023' (i.e. 'left- or right-matches') throughout the calendar year 2022.

Over the course of 2022, 6,730 domain activity events (representing 6,458 unique domain names) were identified for '2023-specific' domains, as shown in Figure 7. 

Figure 7: Daily (top) and monthly (bottom) numbers of new registrations ('N'), re-registrations ('R') and dropped ('D') domains with names beginning or ending with '2023'

Figure 8 shows the growth across 2022 of the cumulative total number of registered domains with names beginning or ending with '2023'.

Figure 8: Daily cumulative total number of registered domains with names beginning or ending with '2023'

Unsurprisingly, the greatest levels of activity (dominated by new registrations) occurred during the latter parts of 2022 (particularly in December), but it is significant that registrations were taking place throughout the year, with a continual growth in the number of registered '2023' domains. It is also worth noting that there were already 2,380 such domains registered at the start of 2022 (compared with 7,524 at the end).

Considering the unique domains represented in the 2022 activity dataset, a range of TLDs (domain extensions) were represented (Figure 9), including significant numbers of new-gTLDs, many of which are of concern due to the previously-noted frequency of their association with infringing activity[12].

Figure 9: Top TLDs amongst the unique '2023' domains represented in the 2022 activity dataset

Significant numbers of these domains were found to be associated with potentially infringing websites, including several with names including top brand names (Figure 10). 

Figure 10: Examples of potentially infringing websites with domain names including references to both '2023' and a brand name from the Interbrand top 20 list of 'best global brands'[13] (SLDs shown in square brackets): (top) potentially fraudulent cryptocurrency-related site [2023-tesla] (registered 27-Dec-2022); (bottom) traffic misdirection / re-direction to a site offering potentially unauthorised or unofficial informational content [2023bmw] and [2023-toyota] (both registered 07-Oct-2022)

A variety of other sites of potential concern were also identified in the dataset, including a range of examples where no brand name was present in the domain name itself. Some of these were, however, found to feature website content which appears to be infringing against specific brands (Figure 11).

Figure 11: Examples of websites offering the sale of potentially counterfeit products and with domain names including a reference to '2023' (SLDs are replicascamisetanba2023 and 2023freerunshoesshop)

Many of the domain names incorporate popular keywords, in apparent attempts to attract traffic in response to common web searches. These included examples such as 'nft' (present in 7 domains) and 'blackfriday' (present in 5 examples, despite Black Friday 2023 being 11 months away). Significantly, 'covid' and 'corona' both appeared in only one example each, perhaps indicating that the online buzz associated with the pandemic is subsiding. The dataset also included some more surprising examples, such as 'keto' (present in 522 domains in the dataset, in addition to several others featuring misspellings such as 'keeto'), perhaps reflective of the continuing popularity of keto diets. Many of these 'keto' domains appear to be part of one or more coordinated registration campaigns, with large numbers of examples with SLDs beginning '2023keto' followed by strings of random characters, across new-gTLDs such as .cyou, .click and .buzz. Even amongst groups of such domains registered on the same day and TLD, a range of content types were observed, including nutrition-related sites, sites advertising a business promotion service provider, and even adult content.

Study 4: Southwest Airlines’ logistics crisis

In December 2022, US air operator Southwest Airlines experienced a 'travel meltdown' in which a series of logistical failures resulted in the cancellation of more than 16,000 flights between 21-Dec and 31-Dec, resulting in tens of thousands of customer refund claims per day, and overall losses to the organisation of between $725 million and $825 million[14]. 

In this case, I considered domains with names containing 'southwest' (or variants), over a one-month period between 12-Dec-2022 and 11-Jan-2023, to determine whether the story generated activity in response to the increased interest in the company and the desire by customers to claim refunds.

Overall, 708 domain activity events, representing 674 unique domain names, were identified during the monitoring period, including a general spike in overall registration activity around the 11-day period in which the incident took place (Figure 12).

Figure 12: Daily numbers of new registrations ('N'), re-registrations ('R') and dropped ('D') domains with names containing 'southwest' (or variants)

Since the term 'southwest' is relatively generic, I then focused on the subset of ('high-relevance') domains which appear relate specifically to Southwest Airlines and the associated events of the story. This was done by considering those domain names which also feature relevant keywords (such as 'air' (but excluding false positives such as 'repairs', 'fairs', etc.), 'aviation', 'aerospace', 'bookings', 'claim' or 'classaction'), or where the domain name itself is a misspelling of Southwest's official website (southwest.com). This yielded a dataset of 46 domain activity events, comprising 43 unique domain names. Within this reduced dataset, the spike in activity around the time of interest can be seen to be much more pronounced (Figure 13).

Figure 13: Daily numbers of new registrations ('N') and re-registrations ('R') combined, and drops ('D'), for high-relevance domains with names containing 'southwest' (or variants)

Of the 36 registration or re-registration events within the dataset of high-relevance domains, 30 (83%) occurred in the four-day period between 27-Dec and 31-Dec.

Of the 43 unique high-relevance domain names in total, 10 were inactive as of the date of analysis (12-Jan-2023). Of the remainder, 27 (68% of the total) resolved to parking pages featuring pay-per-click (PPC) links, indicating an effort by the site owners to monetise the traffic received by the sites. One domain resolved to a site which may be associated with a recruitment scam (Figure 14), one re-directed to the website of a legal-service provider (apparently abusing the Southwest brand name in order to attempt to take advantage of the potential customer desire to take legal action against the company), and one generated a browser warning indicating that dangerous content was formerly present, in addition to other content types.

Figure 14: Example of a website associated with a possible recruitment scam, hosted on a high-relevance, brand-specific domain name

Four (9%) of the high-relevance domain names are configured with MX records, indicating the ability to send and receive e-mails, and suggesting that the domains may be associated with phishing or brand-impersonation activity.

Within the dataset, two instances of domain 'tasting'[15] were identified, comprising domains (with SLDs of southwest-air-line and southwest-bookings) being registered and then dropped the following day, and possibly indicating efforts by the owners to determine the levels of traffic received by the sites, or to launch short-lived (and thereby difficult to detect) phishing attacks.

31 of the high-relevance domains had registration (whois) information available, all of which used privacy-protection providers or had redacted contact information, possibly indicating efforts by the owners to maintain anonymity and potentially nefarious intentions.

Additionally, several individual 'clusters' of domains, potentially representing coordinated registration campaigns by specific entities, were identified. These included:

  • One group of 12 domains all registered on 28-Dec-2022, comprising misspellings of 'southwest.com' and hosted on a group of four consecutive IP addresses
  • One group of five domains all registered on 30-Dec or 31-Dec and all hosted at the same IP address, with names comprising references to 'southwestairlinesclassaction' (or variants)
  • One group of eight domains all registered on 29-Dec or 30-Dec and all hosted at the same IP address

All of the above domains resolved to parking pages featuring PPC links at the time of analysis.

Conclusion

The above news stories or events are all of different types, including examples which are regional or global in scope, and those which may be relevant mainly to specific corporations or industry areas. However, in all cases, resulting spikes in associated domain registration activity were observed. In general, this activity incorporates a mixture of both legitimate and non-legitimate (potentially threatening) registrations, comprising responses both by the official organisations concerned, and by nefarious bad actors.

The findings highlight that, in addition to the construction and maintenance of official domain portfolios by brand owners - and the protection of critical domains using appropriate domain security measures[16,17] - monitoring for third-party activity remains of crucial importance. Particular additional focus must be taken when external events drive increased public interest in associated content, which can result from industry-relevant events, news stories, marketing activity or product releases, corporate changes, and a range of other factors. Accordingly, the monitoring strategy needs to be flexible enough to evolve in response to emerging issues as they develop. Also key to the protection of the brand is a robust enforcement programme incorporating a wide range of approaches, to ensure the swift takedown of damaging infringing content.

It is also striking that so much of the observed activity is carried out so far in advance of the date of the events themselves, showing the significance of proactivity and timeliness in brand protection initiatives, combined with a robust strategy of defensive registrations, to obtain required domains in advance of their registration by wily third parties.

References

[1] https://www.cscdbs.com/en/resources-news/impact-of-covid-on-internet-security/ 

[2] https://www.cscdbs.com/blog/how-to-manage-the-online-effects-of-the-ukraine-war/

[3] https://www.cscdbs.com/en/resources-news/supply-chain-report-form/

[4] https://www.cscdbs.com/blog/euro-2020-part-3-domains-revisited-and-other-channels/

[5] https://www.cscdbs.com/blog/holiday-shopping-events-part-2/

[6] 'The GameStop saga - how online activity and news stories can create feedback loops', Brand Journal, issue no. 21 (April 2021) (internal CSC publication)

[7] https://www.euronews.com/2022/10/14/truss-timeline-key-events-in-three-months-of-political-chaos-in-british-politics

[8] https://www.cnn.com/uk/live-news/uk-prime-minister-announcement-monday-gbr-intl/index.html

[9] https://www.ox.ac.uk/news/2021-01-13-social-media-manipulation-political-actors-industrial-scale-problem-oxford-report

[10] https://www.cscdbs.com/blog/the-highest-threat-tlds-part-1/

[11] https://www.cscdbs.com/blog/the-highest-threat-tlds-part-2/

[12] https://www.cscdbs.com/en/resources-news/threatening-domains-targeting-top-brands/

[13] https://interbrand.com/best-global-brands-2022-download-form/

[14] https://www.cnn.com/travel/article/southwest-airlines-dot-complaints/index.html

[15] https://www.cscdbs.com/blog/patterns-and-trends-in-domain-tasting-of-the-top-10-global-brands/

[16] https://www.linkedin.com/pulse/holistic-brand-fraud-cyber-protection-using-domain-threat-barnett/

[17] https://www.cscdbs.com/en/resources-news/domain-security-report/ (2022)

This article was first published on 31 January 2023 at:

https://www.linkedin.com/pulse/four-new-case-studies-domain-registration-activity-spikes-barnett/

The top GenAI tactics used by counterfeiters and the importance of IP (Original Version)

Generative artificial intelligence ('GenAI') is a general term used to describe automated systems able to generate complex content o...