# Text Mining: Detect Strings: Word Lookup in a Large Corpus of Phrases Using a Large Dictionary

**URL:** <https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665>\
**Category:** Performance\
**Tags:** question\
**Created:** [December 6, 2021, 5:04pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665 "2021-12-06T17:04:13Z")\
**Posts on this page:** 8\
**Page:** 2

<div class="post-metadata">

**Author:** ![Palli](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/palli/32/3380_2.png) [@Palli](https://discourse.julialang.org/u/Palli)\
**Post date:** [December 7, 2021, 5:37pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/21 "2021-12-07T17:37:27Z")

</div>

> [@aourednik](#):
>
> I could but the function needs to handle strings with variable lengths.

If inline strings work much faster (as they should), then a new (relatively simple) hybrid string type should have similar speed-up for variable-length strings. It’s on my radar, not yet implemented, and neither by anyone else that I know of. Such a package would be external to begin with (no performance downside, not just known to all), but could be integrated into Julia later. There’s not no need for Julia 2 (meaning a breaking syntax change) for this, or for most any speed increase I can think of.

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 7, 2021, 5:43pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/22 "2021-12-07T17:43:48Z")

</div>

I’ll have a look tomorrow, I’ve walked home from work now. I might need to spend my eloctrobike money on something 😄

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 8, 2021, 2:35pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/23 "2021-12-08T14:35:26Z")

</div>

I got myself ready to start and I don’t know R so I don’t know exactly what this code does

---

<div class="post-metadata">

**Author:** ![aourednik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aourednik/32/31527_2.png) [@aourednik](https://discourse.julialang.org/u/aourednik)\
**Post date:** [December 10, 2021, 2:32pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/24 "2021-12-10T14:32:53Z")

</div>

I will propose a first draft of a Julia translation. I might need some time to try to get it apporiximatively right

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 10, 2021, 2:37pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/25 "2021-12-10T14:37:04Z")

</div>

or even some pseudo code 🙂

---

<div class="post-metadata">

**Author:** ![aourednik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aourednik/32/31527_2.png) [@aourednik](https://discourse.julialang.org/u/aourednik)\
**Post date:** [December 10, 2021, 5:52pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/26 "2021-12-10T17:52:40Z")

</div>

The real world use case is XML-annotation of [named entities](https://en.wikipedia.org/wiki/Named_entity) in a corpus of texts. There are 2 parts:

## 1. Creating a hashtable

`reference2_ngrams` is probably not necessary, since creating just the hash-table `myPattern` is enough,. However, using such a hashtable is imposed by the [`stringr::str_replace_all(texts,pattern)`](https://stringr.tidyverse.org/reference/str_replace.html) function in R ; perhaps in Julia it would be easier to just find the pattern and prefix `<l>` and postfix `</l>`

```julia-auto
reference2_ngrams = a string array composed of 1 to 4 words per element, e.g. ["LLKSL PJHCV XLAJP CIIIK" , "IPZHD BSUPK" , "VBUBO" , "YYCFO DQWQR ZGKYO LLSZA", ....]
myPattern = a named vecotr of strings of the same lenght as refernence2_ngrams ["LLKSL PJHCV XLAJP CIIIK":"<l>LLKSL PJHCV XLAJP CIIIK</l>", "IPZHD BSUPK":"<l>IPZHD BSUPK</l>", "VBUBO":"<l>VBUBO</l>" ...] # this is basically a hash table of n-grams to find and their replacement 

```

## 2. Replacing all subphrases by their replacements given by the hashtable

Applxiong this to the textdata from the examples further up this Julia discourse topic

```julia-auto
result = string_replace(textdata,maPattern.keys,myPattern.values)

```

---

<div class="post-metadata">

**Author:** ![lawless-m](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/lawless-m/32/30869_2.png) [@lawless-m](https://discourse.julialang.org/u/lawless-m)\
**Post date:** [December 13, 2021, 1:43pm UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/27 "2021-12-13T13:43:01Z")

</div>

My interpretation is

given

```julia
<xml>
 <texts>
  <text>llksl PJHCV xlajp CIIIK</text>
  <text>IPZHD bsupk</text>
 </texts>
</xml>

```

and

```julia
<xml>
 <refs>
  <ref>PJHCV</ref>
  <ref>CIIIK</ref>
  <ref>IPZHD</ref>
 </refs>
</xml>

```

output

```julia
<xml>
 <texts>
  <text>llksl <l>PJHCV</l> xlajp <l>CIIIK</l></text>
  <text><l>IPZHD</l> bsupk</text>
 </texts>
</xml>

```

?

---

<div class="post-metadata">

**Author:** ![aourednik](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/aourednik/32/31527_2.png) [@aourednik](https://discourse.julialang.org/u/aourednik)\
**Post date:** [December 15, 2021, 12:39am UTC](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665/28 "2021-12-15T00:39:19Z")

</div>

It could be, more simply:  
given

```julia
["llksl PJHCV xlajp CIIIK","IPZHD bsupk"‘]

```

and

```julia
refs = ["PJHCV", "CIIIK","IPZHD"]

```

output

```julia
["llksl <l>PJHCV</l> xlajp <l>CIIIK</l>","<l>IPZHD</l> bsupk"]

```

I just have a database column of data and would like to produce a second column with annotated named entities.

Another use case would be:

### input

textfile1.txt:

```txt
llksl PJHCV xlajp CIIIK

```

textfile2.txt:

```txt
IPZHD bsupk

```

### output

xmlfile1.xml

```xml
<xml>
 <metadata>...</metadata>
 <text>
  llksl <l>PJHCV</l> xlajp <l>CIIIK</l>
 </text>
</xml>

```

xmlfile2.xml

```xml
<xml>
 <metadata>...</metadata>
 <text>
  <l>IPZHD</l> bsupk
 </text>
</xml>

```

But the central part is getting the original array of texts annotated with tags

[Previous page](https://discourse.julialang.org/t/text-mining-detect-strings-word-lookup-in-a-large-corpus-of-phrases-using-a-large-dictionary/72665.md?page=1)
