Reusable
Pre-learning materials
Writing code once and reusing it is one of the biggest productivity and quality wins in a Reproducible Analytical Pipeline. Every time we copy-paste a block of code instead of reusing a function, we create another place for a bug to hide, and another place to remember to update when the logic changes.
This section covers four things that helps make great functions:
- Meaningful names
- Do one thing
- Documenting functions
- Consistent style
1. Meaningful names
A good name tells the reader what something represents or does, without them having to read the code in detail. Choosing a good name can be surpringly tricky! When choosing a function name, try to describe what the function does specifically.
- Describe the purpose. Say what the thing does, specifically
- Write for humans - readable, searchable names. e.g
start_dateis much more understandable (and searchable) thand - Be consistent with names used elsewhere in the project, if you already have
load_sales(), useload_costs()rather thanimportCosts()
Compare:
process_data(x)with:
calculate_percentage_change(new, old)The second name tells the reader what’s happening before they’ve even looked inside the function. Avoid vague names like x, tmp, df2 or do_thing when something more specific is easy to write.
If you find yourself adding a comment to explain how a line of code works, that’s often a sign that you could use a better name. Comments are best saved for explaining the why of the code, not how.
👩💻 Your turn
Look at a script you’ve written recently. Find one variable or function name that isn’t self-explanatory. What would you rename it to?
📖 Resources
2. Functions should do one thing
A well-written function does one thing, does it well, and does only that. Aim for functions that:
- Are small
- Do one thing only
- Take as few parameters as possible
Instead of repeating the same calculation in several places:
sales_growth <- (this_year_sales - last_year_sales) / last_year_sales
profit_growth <- (this_year_profit - last_year_profit) / last_year_profitCapture the logic once:
percentage_change <- function(new, old) {
(new - old) / old
}
sales_growth <- percentage_change(this_year_sales, last_year_sales)
profit_growth <- percentage_change(this_year_profit, last_year_profit)sales_growth = (this_year_sales - last_year_sales) / last_year_sales
profit_growth = (this_year_profit - last_year_profit) / last_year_profitCapture the logic once:
def percentage_change(new, old):
return (new - old) / old
sales_growth = percentage_change(this_year_sales, last_year_sales)
profit_growth = percentage_change(this_year_profit, last_year_profit)percentage_change() has one clear job. It doesn’t know or care whether it’s being used for sales, profit, or anything else. That’s what makes it reusable.
When should you extract a function? Duplication isn’t automatically a problem the first time you see it. A useful rule of thumb is the Rule of Three. Once the same logic shows up a third time, that’s a good time to pull the code into a function.
If a function feels cluttered, look for low-level detail that could be pulled out into a reusable helper function. For example, input checks could move into a helper function of it’s own.
# Before: validation is mixed in with the calculation
get_age <- function(age_days, units = "years") {
if (is.character(age_days)) stop("Age must be a number")
if (age_days < 0) stop("Negative age")
if (age_days > 150 * 365) stop("A bit too old?")
if (units == "years") age_days / 365 else age_days
}
# After: validation is pulled out, and the main function is easier to read
get_age <- function(age_days, units = "years") {
check_age(age_days)
if (units == "years") age_days / 365 else age_days
}
check_age <- function(age_days) {
if (is.character(age_days)) stop("Age must be a number")
if (age_days < 0) stop("Negative age")
if (age_days > 150 * 365) stop("A bit too old?")
}# Before: validation is mixed in with the calculation
def get_age(age_days, units="years"):
if isinstance(age_days, str):
raise ValueError("Age must be a number")
if age_days < 0:
raise ValueError("Negative age")
if age_days > 150 * 365:
raise ValueError("A bit too old?")
if units == "years":
return age_days / 365
else:
return age_days
# After: validation is pulled out, and the main function is easier to read
def check_age(age_days):
if isinstance(age_days, str):
raise ValueError("Age must be a number")
if age_days < 0:
raise ValueError("Negative age")
if age_days > 150 * 365:
raise ValueError("A bit too old?")
def get_age(age_days, units="years"):
check_age(age_days)
if units == "years":
return age_days / 365
else:
return age_daysWhere practical, aim for pure functions. These are functions that, given the same inputs, always return the same output, with no side effects (no file access, no printing, no depending on or changing anything outside the function). Pure functions are easier to test and reuse.
👩💻 Your turn
Look back at a script you’ve written. Find a chunk of code that appears more than once, even with small variations.
Sketch out how you’d turn it into a function:
- What would you call it?
- What inputs would it need?
- What would it return?
- Is it a pure function?
Bring your example to the session, we’ll use a few of these as discussion points.
📖 Resources
- Functions — R for Data Science
- Defining Your Own Python Function by Real Python, section on writing clear, single-purpose functions (Python)
3. Document your functions
A function is easy to reuse when the reader can quickly answer three questions, without reading the implementation:
- What does it do?
- What inputs does it need?
- What does it return?
#' Calculate percentage change
#'
#' @param new Numeric. The new (current) value.
#' @param old Numeric. The original (baseline) value.
#'
#' @return Numeric. The proportional change from `old` to `new`.
#'
#' @examples
#' percentage_change(new = 120, old = 100)
percentage_change <- function(new, old) {
(new - old) / old
}The #' comment block above is the R roxygen2 convention. It only generates a help page (?percentage_change) if your function lives inside an R package. If your script is just stand alone, it won’t do anything automatically. It’s still a good habit to write your function docs in this style even in an ordinary script. It forces you to answer the “what/inputs/returns” questions, and it means your docs are ready to go if that code ever does end up in a package.
def percentage_change(new, old):
"""Calculate the proportional change between two values.
Args:
new (float): The new (current) value.
old (float): The original (baseline) value.
Returns:
float: The proportional change from old to new.
Example:
>>> percentage_change(120, 100)
0.2
"""
return (new - old) / oldDocstrings work the same way in one-off scripts and full packages, and are picked up automatically by help().
👩💻 Your turn
Take the function you sketched out in section 2 and write documentation for it. Either a roxygen2-style block in R, or a docstring in Python.
📖 Resources
4. Consistent style
A consistent style makes code easier to scan, understand and modify. It matters more for reusable code: something meant to be shared or reused needs to be readable by people who didn’t write it. Agree conventions with your team, then let a tool handle the mechanical parts.
# Before
my_function=function(x,y){
x+y
}
# After: styler::style_file() applied
my_function <- function(x, y) {
x + y
}{styler} reformats code automatically; {lintr} flags style problems it can’t fix for you.
# Before
def my_function(x,y):
return x+y
# After: black applied
def my_function(x, y):
return x + yblack or ruff format reformat code automatically; ruff also flags issues it can’t fix.
The logic hasn’t changed here, just the formatting. But consistent spacing and layout mean anyone on the team can read this function without first adjusting to a one-off style.
👩💻 Your turn
Run an auto-formatter on a script of your own styler::style_file() (R) or black / ruff format (Python). What did it change automatically? Was anything left for you to fix by hand?
What coding conventions does your team already follow, formally or informally? How do you
You don’t have to run a formatter manually every time. Editors such as VS Code can format your code automatically when you save. For Python, the Ruff VS Code extension can be configured to format on save. For R, Air is a fast R formatter with VS Code and Positron support, and can also format your code on save.
📖 Resources
- {lintr} (R)
- ruff (python)
- Tidyverse Style Guide, sections on functions and syntax (R)
- PEP 8 Style Guide for Python Code, sections on Naming Conventions and Programming Recommendations (Python)
Before the session
Come prepared to discuss:
- What code have you written that would be useful to reuse?
- Share a function you have designed and documented (name, inputs, return value)
- What coding conventions does your team already follow, formally or informally?
- Have you encountered code that was hard to understand or reuse? What made it difficult?