Reusable

Pre-learning materials

Writing code once and reusing it is one of the biggest productivity and quality wins in a Reproducible Analytical Pipeline. Every time we copy-paste a block of code instead of reusing a function, we create another place for a bug to hide, and another place to remember to update when the logic changes.

This section covers four things that helps make great functions:

  1. Meaningful names
  2. Do one thing
  3. Documenting functions
  4. Consistent style
Note

Read through the sections below and try out the 👩‍💻 Your turn sections. Share your answers in the shared learning document for your group. We’ll use your comments as discussion points in the contact session.


1. Meaningful names

A good name tells the reader what something represents or does, without them having to read the code in detail. Choosing a good name can be surpringly tricky! When choosing a function name, try to describe what the function does specifically.

  • Describe the purpose. Say what the thing does, specifically
  • Write for humans - readable, searchable names. e.g start_date is much more understandable (and searchable) than d
  • Be consistent with names used elsewhere in the project, if you already have load_sales(), use load_costs() rather than importCosts()

Compare:

process_data(x)

with:

calculate_percentage_change(new, old)

The second name tells the reader what’s happening before they’ve even looked inside the function. Avoid vague names like x, tmp, df2 or do_thing when something more specific is easy to write.

Tip

If you find yourself adding a comment to explain how a line of code works, that’s often a sign that you could use a better name. Comments are best saved for explaining the why of the code, not how.

👩‍💻 Your turn

Look at a script you’ve written recently. Find one variable or function name that isn’t self-explanatory. What would you rename it to?

📖 Resources


2. Functions should do one thing

A well-written function does one thing, does it well, and does only that. Aim for functions that:

  • Are small
  • Do one thing only
  • Take as few parameters as possible

Instead of repeating the same calculation in several places:

sales_growth <- (this_year_sales - last_year_sales) / last_year_sales
profit_growth <- (this_year_profit - last_year_profit) / last_year_profit

Capture the logic once:

percentage_change <- function(new, old) {
  (new - old) / old
}

sales_growth <- percentage_change(this_year_sales, last_year_sales)
profit_growth <- percentage_change(this_year_profit, last_year_profit)
sales_growth = (this_year_sales - last_year_sales) / last_year_sales
profit_growth = (this_year_profit - last_year_profit) / last_year_profit

Capture the logic once:

def percentage_change(new, old):
    return (new - old) / old

sales_growth = percentage_change(this_year_sales, last_year_sales)
profit_growth = percentage_change(this_year_profit, last_year_profit)

percentage_change() has one clear job. It doesn’t know or care whether it’s being used for sales, profit, or anything else. That’s what makes it reusable.

When should you extract a function? Duplication isn’t automatically a problem the first time you see it. A useful rule of thumb is the Rule of Three. Once the same logic shows up a third time, that’s a good time to pull the code into a function.

If a function feels cluttered, look for low-level detail that could be pulled out into a reusable helper function. For example, input checks could move into a helper function of it’s own.

# Before: validation is mixed in with the calculation
get_age <- function(age_days, units = "years") {
  if (is.character(age_days)) stop("Age must be a number")
  if (age_days < 0) stop("Negative age")
  if (age_days > 150 * 365) stop("A bit too old?")
  if (units == "years") age_days / 365 else age_days
}

# After: validation is pulled out, and the main function is easier to read
get_age <- function(age_days, units = "years") {
  check_age(age_days)
  if (units == "years") age_days / 365 else age_days
}

check_age <- function(age_days) {
  if (is.character(age_days)) stop("Age must be a number")
  if (age_days < 0) stop("Negative age")
  if (age_days > 150 * 365) stop("A bit too old?")
}
# Before: validation is mixed in with the calculation
def get_age(age_days, units="years"):
    if isinstance(age_days, str):
        raise ValueError("Age must be a number")
    if age_days < 0:
        raise ValueError("Negative age")
    if age_days > 150 * 365:
        raise ValueError("A bit too old?")

    if units == "years":
        return age_days / 365
    else:
        return age_days


# After: validation is pulled out, and the main function is easier to read
def check_age(age_days):
    if isinstance(age_days, str):
        raise ValueError("Age must be a number")
    if age_days < 0:
        raise ValueError("Negative age")
    if age_days > 150 * 365:
        raise ValueError("A bit too old?")


def get_age(age_days, units="years"):
    check_age(age_days)

    if units == "years":
        return age_days / 365
    else:
        return age_days

Where practical, aim for pure functions. These are functions that, given the same inputs, always return the same output, with no side effects (no file access, no printing, no depending on or changing anything outside the function). Pure functions are easier to test and reuse.

👩‍💻 Your turn

Look back at a script you’ve written. Find a chunk of code that appears more than once, even with small variations.

Sketch out how you’d turn it into a function:

  • What would you call it?
  • What inputs would it need?
  • What would it return?
  • Is it a pure function?

Bring your example to the session, we’ll use a few of these as discussion points.

📖 Resources


3. Document your functions

A function is easy to reuse when the reader can quickly answer three questions, without reading the implementation:

  • What does it do?
  • What inputs does it need?
  • What does it return?
#' Calculate percentage change
#'
#' @param new Numeric. The new (current) value.
#' @param old Numeric. The original (baseline) value.
#'
#' @return Numeric. The proportional change from `old` to `new`.
#'
#' @examples
#' percentage_change(new = 120, old = 100)
percentage_change <- function(new, old) {
  (new - old) / old
}
Important

The #' comment block above is the R roxygen2 convention. It only generates a help page (?percentage_change) if your function lives inside an R package. If your script is just stand alone, it won’t do anything automatically. It’s still a good habit to write your function docs in this style even in an ordinary script. It forces you to answer the “what/inputs/returns” questions, and it means your docs are ready to go if that code ever does end up in a package.

def percentage_change(new, old):
    """Calculate the proportional change between two values.

    Args:
        new (float): The new (current) value.
        old (float): The original (baseline) value.

    Returns:
        float: The proportional change from old to new.

    Example:
        >>> percentage_change(120, 100)
        0.2
    """
    return (new - old) / old

Docstrings work the same way in one-off scripts and full packages, and are picked up automatically by help().

👩‍💻 Your turn

Take the function you sketched out in section 2 and write documentation for it. Either a roxygen2-style block in R, or a docstring in Python.

📖 Resources


4. Consistent style

A consistent style makes code easier to scan, understand and modify. It matters more for reusable code: something meant to be shared or reused needs to be readable by people who didn’t write it. Agree conventions with your team, then let a tool handle the mechanical parts.

# Before
my_function=function(x,y){
x+y
}

# After: styler::style_file() applied
my_function <- function(x, y) {
  x + y
}

{styler} reformats code automatically; {lintr} flags style problems it can’t fix for you.

# Before
def my_function(x,y):
    return x+y

# After: black applied
def my_function(x, y):
    return x + y

black or ruff format reformat code automatically; ruff also flags issues it can’t fix.

The logic hasn’t changed here, just the formatting. But consistent spacing and layout mean anyone on the team can read this function without first adjusting to a one-off style.

👩‍💻 Your turn

Run an auto-formatter on a script of your own styler::style_file() (R) or black / ruff format (Python). What did it change automatically? Was anything left for you to fix by hand?

What coding conventions does your team already follow, formally or informally? How do you

Tip

You don’t have to run a formatter manually every time. Editors such as VS Code can format your code automatically when you save. For Python, the Ruff VS Code extension can be configured to format on save. For R, Air is a fast R formatter with VS Code and Positron support, and can also format your code on save.

📖 Resources


Before the session

Come prepared to discuss:

  • What code have you written that would be useful to reuse?
  • Share a function you have designed and documented (name, inputs, return value)
  • What coding conventions does your team already follow, formally or informally?
  • Have you encountered code that was hard to understand or reuse? What made it difficult?