During the analysis of source content by a CAT tools, words are counted and grouped into three categories of words: matches, fuzzy matches, and new words.
Each of these word categories is further subdivided into related sub-categories. Discussion of the meaning of each word count category follows.
100% match
A 100% match is when a new segment for translation is already contained in the TM. For instance, let’s say the client’s TM contains the phrase, “My favorite color is red.” Then, the client sends new content for translation that also contains the phrase, “My favorite color is red.” In theory, that segment should not need to be translated again. With a well-maintained TM the translation already contained therein would simply be applied.
However, many TMs on the market are not well maintained. Issue-TMs may contain multiple translations for a single segment as the result of rush work assigned out to multiple translators. When standards are not established for how content should be translated, human variation naturally occurs. This is not so problematic with a message as simple as “My favorite color is red.” However, variation in the application of key subject field terminology may impede target user experience.
The suitability of a 100% TM match leveraged into new translation requests from clients should be checked no matter if the TM is considered well maintained or not. Context changes meaning to greater and lesser degrees and may require that the applied translation be updated to account for number and gender, such as when the abbreviate of “each” used in counting is abbreviated as “ea.” which is then interpreted as the CAT tool as the end of the segment do to a full stop period. In this case, cada uno may need to be updated to cada una or vice versa, such as in the below example of the translation from a price list, in which the phrase “$1.50 ea.” is a 100% match in English and is not a 100% match in Spanish (unless translated as c/u).
| Source | Target |
| Apples | Manzanas |
| $1.50 ea. | $1,50 cada una |
| Plantains | Platanós |
| $1.50 ea. | $1,50 cada uno |
Some practitioners and providers have informal policies of not reviewing 100% TM matches, however, to accommodate rush delivery requests. In the case of a price list, the use of “$1,50 cada una” where “$1,50 cada uno” should be used is only a little embarrassing. In the case of kitchen terminology in Czech, the incorrect application of two potential translations for cook (one word for “cooking on the stove” and another word for “cooking in the oven”) could mean a failed recipe. As the content gets riskier, such as in medicine, mechanics, and law, the impact of the application of false positive 100% matches in translations gets higher and costlier to fix with damages sometimes being irrecoverable, such as in cases of death.
101% match
Having defined 100% matches, we can now define 101% matches. A 101% match is better contextual match because larger chunks of segments are grouped together in the same way in both the TM and in any content submitted for translation. Returning to our example of “each,” let’s say the client submits another price list for translation. The price contains some new content and some content that has already been translated. Our TM already contains the translations of the segments Apples, $1.50 ea., Bananas, $1.50 ea., in the same order as the segments from the content submitted for translation. The middle two of those segments, $1,50 cada una and Platanós, are 101% matches since the immediate segment that proceeds and comes after each of these matches is exactly the same. A 101% match therefore represents a greater level of reliability in the case of well-maintained TMs.
| Source | Target |
| Apples | Manzanas |
| $1.50 ea. | $1,50 cada una |
| Bananas | Platanós |
| $1.50 ea. | $1,50 cada uno |
| Flowers | Flores |
| $1.50 ea. | $1,50 cada una |
Repetitions
A repetition is a segment that is exactly the same as one that has already appeared in content for translation. That is, the same segment appears over and over again. Generally, CAT tools are setup so that the translation of the first instance of the repetition is automatically applied to all other instances of a repetition. This automation speeds translation though all instances of auto-propagation should be checked for contexts in which the translation of the repetition is a false match. In practice, repetitions are often first to be overlooked when rush deadlines are applied. This is a risk LPMs need to be aware of when assigning rush projects.
Fuzzy matches
A fuzzy match is a segment for which a partial match exists in a TM. In CAT analysis, fuzzy matches are presented according to percentage rates which are determined based on the percentage of words in a segment from content for translation that match existing words in the segments in the TM.
For example, let’s say our TM for our client contains a translation of the phrase, “My favorite color is red.” Then the client submits new content for translation, which contains the phrase, “My favorite color is blue.” It’s easy to see that this new content for translation would constitute around an 80% match since 4 of the 5 words in both sentences are the same.
Negotiation is needed to establish the thresholds for fuzzy matches. In general, fuzzy matches are defined as 75-99% matches, but providers choose their own thresholds. Some may define fuzzy matches as 85-99% matches, in which case the translation would cost more. Some define fuzzy matches all the way down from 50-99% matches. It may be tempting to define fuzzy matches all the way down to 50% to decrease translation costs, but this definition also impacts the speed with which a translation can be delivered. It’s often faster to start from scratch than try to revise really fuzzy matches that simply don’t apply to the context. A translator revising 50% fuzzy matches may take longer than if they had just translated that content from scratch.
Internal leveraging
Internal leveraging is among the newest of the categories of word types from CAT analysis. In internal leveraging, content for translation is analyzed against itself (not against the TM). The three sentences that follow are examples of segments that are 80% matches amongst themselves without needing to analyze the content against a TM: My favorite color is red. My favorite color is blue. My favorite color is yellow. Internal leveraging can be an important indicator of speed, since the first instance of an internal fuzzy match can be auto-propagated into later instances.
When working with translators, LPMs should be aware of whether internal leveraging is being applied to translation rates or not. Translators often view increased automation as a threat to their rates. The word counts generated by different CAT tools for a single document is already varied, since each CAT tool uses its own algorithm to product word counts. This can be a source of major discrepancy if a translator verifies your word counts with a different CAT tool. When word count discrepancies are heightened because an agency has applied internal fuzzies within the pay range for fuzzy matches without alerting the translator, misunderstandings about quality expectations may result, following the “you get what you pay for” adage.
New words
New words are a bit of a misnomer in that according to the term one would expect a new word to be a word for which no match exists. This category of new words also tends to include at least 0-50% matches, though all the way up to 84% matches are sometimes defined as new words. Each provider defines their own expectations on these thresholds.
Leave a Reply