# Reading word '.doc' file

**URL:** <https://discourse.julialang.org/t/reading-word-doc-file/16506>\
**Category:** New to Julia\
**Tags:** question\
**Created:** [October 18, 2018, 6:29pm UTC](https://discourse.julialang.org/t/reading-word-doc-file/16506 "2018-10-18T18:29:13Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![hasanOryx](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/hasanoryx/32/5373_2.png) [@hasanOryx](https://discourse.julialang.org/u/hasanOryx)\
**Post date:** [October 18, 2018, 6:29pm UTC](https://discourse.julialang.org/t/reading-word-doc-file/16506/1 "2018-10-18T18:29:13Z")

</div>

Hi,  
How can I read the content of MS word ‘.dic’ file, and search for key word in it.  
I’d like to make quick scan for some CVs.

Thanks

---

<div class="post-metadata">

**Author:** ![mgiugliano](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/mgiugliano/32/5818_2.png) [@mgiugliano](https://discourse.julialang.org/u/mgiugliano)\
**Post date:** [October 18, 2018, 9:44pm UTC](https://discourse.julialang.org/t/reading-word-doc-file/16506/2 "2018-10-18T21:44:09Z")

</div>

The easiest way would involve the use of an external utility (e.g. PanDoc) to convert a Word .doc file into a text file (see [https://gist.github.com/aembleton/1eb889bc443996a508df](https://gist.github.com/aembleton/1eb889bc443996a508df)). Once such a preprocessing has been completed, it should be straightforward to load the text file from Julia and parse/search it.

---

<div class="post-metadata">

**Author:** ![tkoolen](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tkoolen/32/1603_2.png) [@tkoolen](https://discourse.julialang.org/u/tkoolen)\
**Post date:** [October 19, 2018, 3:19am UTC](https://discourse.julialang.org/t/reading-word-doc-file/16506/3 "2018-10-19T03:19:26Z")

</div>

Could maybe use the Python [docx](https://python-docx.readthedocs.io/en/latest/) package (if you mean docx instead of doc) via PyCall.
