# Large programs: structuring modules & include such that to increase performance and readability

**URL:** <https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102>\
**Category:** Performance\
**Tags:** question, modules, code-organization\
**Created:** [September 24, 2019, 9:34am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102 "2019-09-24T09:34:42Z")\
**Posts on this page:** 15\
**Page:** 1

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [September 24, 2019, 9:34am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/1 "2019-09-24T09:34:43Z")

</div>

> I am writing a large hydrological model for which I require to divide the program in many _modules_ incorporated in individual files. I am questioning what will be an efficient and readable approach to share the modules across the numerous files?
> 
> I found that for e.g. `Include(“Filename.jl”)` in every module is clean but not the fastest especially when I need to perform loops (which will call _include_ many times.)
> 
> I found that an **efficient** way is to ` include(FileName)` in the **MAIN()** program and then to export modules as variables.

` eg. EVAPOTRANSPIRATION(, evaporation, transpiration, ...)`

> It works but it is not elegant, so I am questioning if there is a more elegant way of sharing modules?

Below is an example which works:

**FILE: Main.jl**

```julia
 module main
      include(“Evaporation.jl”)
      include(“Transpiration.jl”)
     include("EvapoTranspiration.jl")
 
      function MAIN(Et)
           for i =1:10
                return EvapoTranspiration = evapoTranspiration.EVAPOTRANSPIRATION(Et * i, evaporation, transpiration)
           end
      end
end

```

**FILE: EvapoTranspiration.jl**

```julia
module evapoTranspiration
     export EVAPOTRANSPIRATION

     function EVAPOTRANSPIRATION(Et, evaporation, transpiration)
          EvapoTranspiration = evaporation.EVAPORATION(Et)+transpiration.TRANSPIRATION(i * Et)
     end 
end

```

**File: Evaporation.jl**

```julia
module evaporation
     export EVAPORATION

     function EVAPORATION(Et)
          return Evaporation = Et * 0.3
    end
end

```

**File: Transpiration.jl**

```julia
module transpiration
     export TRANSPIRATION

     function TRANSPIRATION(Et)
          return Transpiration = ET * 0.5
     end
end

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 24, 2019, 10:23am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/2 "2019-09-24T10:23:27Z")

</div>

> [@JosephPollacco](#):
>
> `Include(“Filename.jl”)` in every module is clean but not the fastest especially when I need to perform loops (which will call _include_ many times.)

Generally, there is no reason to call `include` anywhere but the toplevel. You definitely should not `include` files in a loop over and over. There is no point in it either, definitions in those files should be available once they are included.

It is not clear from your code whether you want/need [submodules](https://docs.julialang.org/en/v1/manual/modules/#Relative-and-absolute-module-paths-1), but that’s what you effectively get above.

Whether you need submodules or multiple packages is up to you, I would say that around 5–10k LOC is the threshold for most people (and that’s a _lot_ of code in Julia).

I would suggest partitioning the code into a _package_ which contains reusable functionality, and _scripts_, which are `using` this package and run things.

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [September 24, 2019, 10:17pm UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/3 "2019-09-24T22:17:50Z")

</div>

> [@JosephPollacco](#):
>
> evaporation, transpiration

Dear Tamas I thank you for responding to my question. I agree that one should `call include` in the _toplevel_ which in this example is in `module main`.

**To my understanding the question you asked is to determine if we put all the modules in one files and use the submodules architecture.** Since we are dealing with large models with different components of the water cycle it will be not readable to write one big file such that in this e.g. `Main.jl` will include all the modules`evaporation, transpiration`. into one big file. To my understanding it will be best to partition the different processes into different files

**You suggested to create packages.** I agree that different tools should be partitioned into _packages_, but to my understanding I do not think that the scientific modules should be put into packages which will decreases the visibility and the easiness of correction of the code.

**The question remains, what is a clean way of making the functions of the modules declared at the top level available to all modules in the program written in different files?**

---

<div class="post-metadata">

**Author:** ![ffevotte](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ffevotte/32/6587_2.png) [@ffevotte](https://discourse.julialang.org/u/ffevotte)\
**Post date:** [September 24, 2019, 10:56pm UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/4 "2019-09-24T22:56:42Z")

</div>

I think you are asking two relatively different (and largely orthogonal) questions:

1. “physical” organization: should the source code be split into several files ?
2. “logical” structure: should the code be structured into (sub-)modules ?

I would say that the answer to question 1. is most certainly “yes”. As soon as your code starts growing, you’re probably better off splitting it into several source files. This is what `include()` is for.

This is largely orthogonal to modules. You can for example start with a unique source file defining the top-level module of your project. When this file grows too big for your taste, you split it into several files and include these in the main file. The `include`d files do not have to define new sub-modules; they may as well contain the exact same content that originally was copied from the main source file.

As for question 2, you might also want to “logically” structure the code into (sub-)modules. Sub-modules help you group related features together, so that some client code can issue a `using SubModule` and get access to everything that is exported by the module. Again, this “logical” structure is largely orthogonal to the “physical” organization of your sources into files : it would for example be perfectly legal to have a unique source file defining the top-level module and all sub-modules. Nevertheless, it is customary (and good practice) to define each sub-module in a source file of the same name (which can then `include` other source files if needed.)

Now a third question would be to ask whether sub-modules defined in your project should rather be full-fledged packages. This would be the case if such sub-modules provide features that could be useful in several contexts. In this case, the source code for the sub-module should be put into a different package. And this package shoud be `Pkg.add`ed as a dependency in your project.

  

Is this clear? If you can briefly describe the various components in your project, we might be able to help you determine an adequate way of organizing and structuring you source code.

---

<div class="post-metadata">

**Author:** ![ffevotte](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ffevotte/32/6587_2.png) [@ffevotte](https://discourse.julialang.org/u/ffevotte)\
**Post date:** [September 24, 2019, 11:11pm UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/5 "2019-09-24T23:11:22Z")

</div>

> [@JosephPollacco](#):
>
> The question remains, what is a clean way of making the functions of the modules declared at the top level available to all modules in the program written in different files?

To answer this specific question: if you want to define submodules (and, again, it remains unclear whether submodules are useful/needed in your case), then each part of the code wanting to use functions defined in another submodule should issue a `using SubModule` statement.

For example:

Main.jl

```julia
module Main
    # equivalent to copy-pasting the code contained in Evaporation.jl
    # since the code in question defines a module called Evaporation,
    # this will become a submodule Main.Evaporation
    include("Evaporation.jl")

    # Same as above
    include("EvapoTranspiration.jl")

    # This brings evapo_transpiration into scope
    using .EvapoTranspiration
    
    main() = println(evapo_transpiration(42.))
end

```

  

Evaporation.jl

```julia
module Evaporation
    # If some client code issues a `using Evaporation` statement,
    # the function `evaporation` will be brought into scope
    export evaporation

    # Actually define the function
    evaporation(x) = x
end

```

  

EvapoTranspiration.jl

```julia
module EvapoTranspiration
    export evapo_transpiration

    # Use a relative path to refer to the Evaporation module:
    # two leading dots mean that the module is defined as a submodule
    # of the current parent
    using ..Evaporation

    # the `evaporation` can be used directly, since the Evaporation module
    # has been brought into scope
    evapo_transpiration(x) = 2 * evaporation(x)
end

```

---

<div class="post-metadata">

**Author:** ![Tamas\_Papp](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tamas_papp/32/25949_2.png) [@Tamas\_Papp](https://discourse.julialang.org/u/Tamas_Papp)\
**Post date:** [September 25, 2019, 6:06am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/6 "2019-09-25T06:06:25Z")

</div>

@ffevotte already answered your question about code organization: just split into files, and include them. For example, [this is a typical layout](https://github.com/tpapp/DynamicHMC.jl/blob/ace360a137adfb9fa3c3cf291910ae21c4f68fda/src/DynamicHMC.jl) I usually use.

Scientific code should also be put into packages (except for the runtime code). Julia packages are so lightweight that it doesn’t take much, and it will make your life much easier, eg you can use

> **[GitHub - timholy/Revise.jl: Automatically update function definitions in a...](https://github.com/timholy/Revise.jl)**
>
> Automatically update function definitions in a running Julia session - GitHub - timholy/Revise.jl: Automatically update function definitions in a running Julia session

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [September 27, 2019, 7:39am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/7 "2019-09-27T07:39:18Z")

</div>

> [@ffevotte](#):
>
> two leading dots mean that the module is defined as a submodule # of the current parent using …Evaporation

Thanks ffevotte, you answered my question. Thanks to your help I have cleaned up my code and it works beautifully.

Your answer _that two leading dots mean that the module is defined as a submodule of the current parent using …Evaporation_

The question why we use 2 … and not 1 . ?

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [September 27, 2019, 7:55am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/8 "2019-09-27T07:55:22Z")

</div>

Thanks I have included Revise.jl in the main module.

---

<div class="post-metadata">

**Author:** ![Raf](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/raf/32/3383_2.png) [@Raf](https://discourse.julialang.org/u/Raf)\
**Post date:** [September 27, 2019, 8:33am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/9 "2019-09-27T08:33:29Z")

</div>

Great to see more biophysical models in julia! We should start an organisation soon to get everyone talking more.

And I have some recomendations.

In julia structure doesn’t affect performance much if at all. It all compiles together the same way whether you have separate packages or one repository. You should consider the implications of that for collaborating and sharing code in these models, it’s _much_ easier than with Fortran and C/C++ so leveraging the community is nearly always the best strategy - for math tools like DifferentialEquations.jl, optimisers and quadratic solvers that often seem to be written custom in Fortran models. It also means we can work towards sharing other components - like [photosynthesis](https://github.com/rafaqz/Photosynthesis.jl). It has really no cost using external packages.

If you must use one large repository I would suggest using internal modules and make them as modular as possible like they are doing in [CLIMA](https://github.com/climate-machine/CLIMA). Then your program structure stays more manageable and you can always separate out modules to a separate package if external use-cases emerge.

But I increasingly make separate packages for everything as I end up using modelling components in multiple projects, which means other people can use them too.

Lastly the style guide has some general patterns and following them really helps readability among other julia users [Style Guide · The Julia Language](https://docs.julialang.org/en/v1/manual/style-guide/index.html)

---

<div class="post-metadata">

**Author:** ![ffevotte](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/ffevotte/32/6587_2.png) [@ffevotte](https://discourse.julialang.org/u/ffevotte)\
**Post date:** [September 27, 2019, 9:03am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/10 "2019-09-27T09:03:42Z")

</div>

Everything is explained in this part of the documentation:

[https://docs.julialang.org/en/v1/manual/modules/#Relative-and-absolute-module-paths-1](https://docs.julialang.org/en/v1/manual/modules/#Relative-and-absolute-module-paths-1)

The simplest way I like to think of it is:

- not dot means that the module is identified by an absolute path:

- the first `.` switches from an absolute path to a relative one:

- every additional leading `.` goes up one level in the modules hierarchy:

Does that make sense?

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [September 28, 2019, 7:06am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/11 "2019-09-28T07:06:03Z")

</div>

Thanks Ffevotte for your detailed and useful explanations, it makes perfectly sense .😀

---

<div class="post-metadata">

**Author:** ![foobar\_lv2](https://avatars.discourse-cdn.com/v4/letter/f/ee59a6/32.png) [@foobar\_lv2](https://discourse.julialang.org/u/foobar_lv2)\
**Post date:** [September 28, 2019, 11:52am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/12 "2019-09-28T11:52:23Z")

</div>

I would definitely recommend submodules instead of plain `include` if you plan to have different people working on them. Modules in julia serve to separate namespaces, and you don’t want name-clashes between helper functions or constants like `_compute_stuff` in evaporation and transpiration. Furthermore, this simplifies debugging.

Regarding exports and `using`, this is pure syntactic sugar once fully qualified names start hurting readability (`Evaporation.evaporate(...)` is much more readable if you have few call-sites, because the reader doesn’t need to look up what module this comes from, but the verbosity hurts readability if call-sites are all over the place; I suggest always starting with plain `import`, and refactoring to `using` once you have many call-sites).

Regarding packaging and physical organization, think about whether it will be common that a single git commit needs to touch several of your modules. It sounds like that will be the case; hence I would recommend sticking to a single git repo and a single package for the beginning. Regarding physical organization into several source files, the same applies: Too many tiny files hurt readability, too few huge files cause perpetual merge conflicts in git.

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [October 3, 2019, 10:10pm UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/13 "2019-10-03T22:10:02Z")

</div>

Thanks for your comments which is really helpfull

---

<div class="post-metadata">

**Author:** ![JosephPollacco](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/josephpollacco/32/51404_2.png) [@JosephPollacco](https://discourse.julialang.org/u/JosephPollacco)\
**Post date:** [October 3, 2019, 10:11pm UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/14 "2019-10-03T22:11:35Z")

</div>

Thanks the style guide is helpfull and your recommendations are usefull

---

<div class="post-metadata">

**Author:** ![Tero\_Frondelius](https://sea2.discourse-cdn.com/julialang/user_avatar/discourse.julialang.org/tero_frondelius/32/7629_2.png) [@Tero\_Frondelius](https://discourse.julialang.org/u/Tero_Frondelius)\
**Post date:** [October 4, 2019, 9:43am UTC](https://discourse.julialang.org/t/large-programs-structuring-modules-include-such-that-to-increase-performance-and-readability/29102/15 "2019-10-04T09:43:59Z")

</div>

Please, can you add your dot explanation to Julia documentation? The last time I checked, it wasn’t as well explained.
