From Problem Statement to Defensible Model: The First Two Days of a HiMCM Paper

From Problem Statement to Defensible Model: The First Two Days of a HiMCM Paper

In a fixed order. Strip the problem statement into a numbered requirements checklist, restate the question you will actually answer, write an assumption ledger where every assumption carries a justification, define your variables and units, then choose the simplest model family that can produce the required output. What earns credit is that trail of decisions, not a number.

Teams that have trained on AMC-style questions find the first hours of HiMCM disorienting, and the reason is structural rather than mathematical. There is nothing to look up, nothing to check, and no moment where a number tells you that you are right. COMAP describes the contest as an exercise in modelling, problem solving and writing, and the judging rewards a complete, logical, organised argument rather than a correct answer — which is why HiMCM is not a harder AMC. This article is about the part almost nobody drills: the disciplined conversion of a paragraph of English into a mathematical object you can defend.

Six-stage pipeline from HiMCM problem statement to model: number the requirements, restate the question, write the assumption ledger, define variables and units, choose the simplest sufficient model family, then run a baseline and one extension, with each stage producing a named artefact that appears in the paper.
The order matters more than the mathematics: each stage produces something that ends up in the submitted paper. Section list per COMAP's published guidance on what the best papers contain.

Stage one: turn the prompt into a numbered requirements checklist

COMAP releases the problems at 3pm EST on the first day of the window, and HiMCM teams choose between Problem A and Problem B. Its own advice is blunt: read all the available problem statements and requirements carefully before choosing, because each problem has different and specific requirements, and submit a solution to only one problem. That word requirements is the hinge. A HiMCM problem is not one question; it is a short brief containing a numbered or bulleted list of things the client wants.

So the first artefact is mechanical. Copy every explicit ask into a table, one row each, and label them R1, R2, R3. Beside each, write the form of the answer: a single number, a ranked list, a map, a table of scenarios, a recommendation letter. Then, at the end of day thirteen, you check the paper against that table. Papers lose tiers for a boring reason far more often than for weak mathematics: they answer R1 and R2 beautifully and forget R4 entirely.

Two details worth flagging while you build the checklist. If a problem asks for a letter or memo to a decision-maker, that item is a deliverable in its own right and COMAP is specific about anonymity — no names anywhere in the paper, and if you need a closing, sign as your team number rather than as yourselves. And if the requirement mentions data, check the contest site: COMAP states that any relevant data files or supporting materials required for its problems are included there.

Stage two and three: restate the question, then earn every assumption

Before modelling, write one sentence of the form: we will estimate X, measured in Y, for the population Z, over horizon T. If four people cannot agree on that sentence, you do not yet have a shared problem, and the fastest way to waste day three is to start coding while that disagreement is unresolved. The restatement is also the sentence the summary sheet is built around.

Then comes the ledger. COMAP states that the best papers usually include sections addressing assumptions with justifications, and that pairing is the whole point: an assumption without a justification is not simplification, it is an unexplained liability a judge will find. Because you may not consult anyone outside your team during the window, a justification has to rest either on a source you cite or on an argument a reader can follow. Four columns are enough.

Assumption Justification a judge can check What it buys the model How it could fail
Demand is constant within each hour Cited hourly data are reported as hourly averages; sub-hourly variation is not published Lets us treat arrivals as a rate, not a schedule Underestimates short peaks; tested in sensitivity
No growth in population over the horizon Horizon is two years; cited growth rate is below one per cent Removes one time-varying parameter Fails for a longer horizon; we state the limit
All units behave identically Published specifications differ by under five per cent Allows a single representative parameter Hides tail behaviour; noted in weaknesses
Costs scale linearly with quantity Price list is per unit with no listed volume break Keeps the objective function linear and solvable Breaks if discounts exist; extension tests a step cost
Our own illustrative ledger format, not a COMAP template. The pattern to copy is the second column: an assumption with a checkable reason attached.

Notice what the fourth column does. It writes your strengths-and-weaknesses section and your sensitivity plan for you, in advance, while you still have days left to act on them. Teams that skip the ledger end up inventing weaknesses on day thirteen, and it shows: the limitations read like an apology instead of an analysis.

Stage five: choose the simplest model family that can produce the required output

Most teams choose a technique because they know it, then bend the problem to fit. Work in the other direction. Look at the verb in each requirement, list the families that could produce that output, pick the simplest one that can, and say in the paper which alternative you rejected and why. That single sentence — we considered a stochastic queueing model but the data available support only a deterministic peak-load estimate — is worth more than an extra model nobody validated.

What the requirement asks for Families worth considering first What you must then justify
How many, in future years Growth curve, regression on cited series, difference equation Why that functional form, and the horizon you trust
Where to place or locate something Optimisation, facility location, clustering on a distance metric The objective function and the metric behind it
Which option is best, or a ranking Weighted scoring, multi-criteria comparison, dominance analysis Where the weights come from, plus weight sensitivity
How something spreads or evolves Compartmental model, difference equations, agent or event simulation The transition rules and the time step
How to schedule or allocate under limits Integer or linear programme, greedy heuristic with a bound Constraints, feasibility, and how close to optimal you are
How likely, or how risky Probability model, Monte Carlo over cited parameter ranges Distributions chosen, and how many runs are enough
Our own editorial mapping for triage in the first two days, not a COMAP classification. The right-hand column is where credit is actually won.

Then build the baseline before the sophistication. A working simple model on day four gives you numbers to interrogate for ten days; an ambitious half-built model on day eleven gives you nothing to write about. Add exactly one extension — relax the strongest assumption in your ledger — and compare it with the baseline. That comparison is itself a result, and it is the kind of result judges quote in commentaries.

The justification trail is the deliverable

Three facts about submission change how you should build the model, and the third one decides your tooling before you write any code: per COMAP's contest instructions, a team that uses AI tools must follow COMAP's AI-use policy and append a separate report on that use to the end of its PDF, outside the page limit — confirm the current wording on comap.org. COMAP states that the entire solution goes in as one PDF of at most 25 pages, in English, in at least 12-point type, and that teams should not include programs, software or databases, because they are not used in judging. So the model exists, for judging purposes, only as far as it is legible in prose, equations, pseudocode, tables and figures inside those pages. A brilliant script nobody can read is worth nothing; a clearly specified model with modest mathematics is worth a great deal.

That is also why every modelling decision needs a visible reason. Judges are reading for a chain: this is the question, these are the assumptions and why they are acceptable, this is the model and why this family, these are the results, here is where they are fragile, here is what we would do next. Any link you leave implicit, a judge has to guess — and guessing is exactly what a defensible paper removes. The habit is the same one described in our foundation piece on why modelling is not problem solving: you are not proving a result, you are making a case.

Anatomy of one defensible modelling decision: state the decision, name the alternative rejected, give a checkable justification, then test the decision in sensitivity analysis, with the result feeding back into the strengths and weaknesses section.
One modelling decision, fully documented. Our own editorial framework; the section names follow COMAP's description of what strong papers contain.

A short worked pass, so the sequence is concrete

Take an invented brief of the kind schools use for practice — not a past contest problem: a campus wants to know how many electric-vehicle charging points to install over the next two years, and where. The requirements checklist splits immediately into R1, a number; R2, locations; R3, a justification of the trade-off between cost and waiting time; R4, a one-page memo to the facilities director.

The restatement fixes the ambiguity: we will estimate the minimum number of charging points such that the average wait in the busiest hour stays under fifteen minutes, for staff and visitor vehicles, over a 24-month horizon. Note what that sentence did — it turned a vague ask into a threshold, a population and a horizon, all three of which are now assumptions you must justify rather than choices you smuggled in.

From there the model family follows from R1 and R3: demand plus service time plus a wait threshold points to a queueing or discrete-event simulation, but if the only data you can cite are daily totals, the honest choice is a deterministic peak-load calculation with a stated safety margin, plus a sensitivity sweep over the peak factor. R2 becomes a small placement problem with an explicit objective, such as minimising the mean walking distance from the parking bays you can identify. R4 is written last and signed as your team number. Total mathematics involved: a rate calculation, a bounded optimisation, and a sweep. None of it advanced — all of it defensible, which is the standard the contest actually applies.

Frequently asked questions

Do we have to use advanced mathematics to win a HiMCM award?
No. COMAP judges the logic of the modelling and the write-up, and accepts partial solutions. A justified simple model beats a complex one.

How many assumptions should a HiMCM paper list?
There is no COMAP number. Our editorial guidance is to list only assumptions that change results, and to attach a checkable justification to each one.

Can we ask a teacher whether our model choice is sensible?
Not during the window. COMAP prohibits seeking ideas or information from anyone outside the team, including the advisor and subject experts.

Should we submit our code with the paper?
No. COMAP states programs, software and databases should not be sent and are not used in judging, so the model must be legible inside the PDF itself.

This is an independent guide operated by Hanlin Education for China-based international-school students. We are not affiliated with, endorsed by, or sponsored by COMAP, which is the sole official authority for HiMCM rules, dates, registration and judging. Confirm all current details on comap.org before acting on anything here. Errors reported to us are corrected within 7 working days.