Housing Data and Technology Innovation Challenge Participant Handbook

This Handbook is to be used in conjunction with the Housing Data and Technology Innovation Challenge Official Rules, available at Official Rules. In the event of any conflict between the Official Rules and this Handbook, the terms of the Official Rules shall prevail.

The Challenge

We invite Teams of eligible Participants (as defined in the Challenge’s Official Rules) to build tools that help decisionmakers understand how state-level housing policy reforms are likely to affect housing production or the aggregate stock relative to a baseline.

Each tool should:

  1. Address housing supply at the state level. Tools should estimate aggregate, statewide impacts of policy reforms on housing production or stock relative to a baseline. Tools may estimate impacts at smaller levels of geography but should aggregate them to the state level.
  2. Respond to a standard set of policy questions. All Submissions must estimate a baseline scenario, which models housing production under the current regulatory and economic environment without policy changes, and also demonstrate how the tool addresses at least three of the five benchmark scenarios (see Benchmark Policy Scenarios , below). This ensures cross-team comparability while allowing teams to focus on the strengths of their approach.
  3. Be transparent, replicable, and explainable. A policymaker without a technical background should understand what your tool does, what it assumes, and what it cannot tell them. Teams may make modeling and policy design assumptions provided those assumptions are clearly described and justified.
  4. Be open source. All Submissions must be published under an approved OSI-compliant license as specified in the Official Rules.

The Challenge is methodology-agnostic. Your tool might be an econometric model, a simplified pro forma simulator, a rule-based policy engine, a geospatial analysis, a data-synthesis dashboard, or something we have not imagined. What matters is that it rigorously and transparently addresses real policy questions and is usable by its intended audience.

Precise forecasts are less important than tools that convey the likely direction, magnitude, uncertainty, and drivers of policy impacts. Successful Submissions will prioritize transparency, interpretability, and practical policy insight over forecasting complexity.

What to Avoid

  • Parcel-level precision. State- or region-level breadth and policy relevance will be prioritized over parcel-level granularity. Tools that sacrifice transparency for fine-grained precision will score lower.
  • Opaque-box models. Machine learning approaches are welcome, but only if they include meaningful interpretability and explainability components. A model that produces numbers without explaining why will not score well on Transparency.
  • Closed or paywalled dependencies without a free-access path. Submissions must be open source. Tools that rely on paid licenses, proprietary APIs, or closed-source dependencies are eligible only if there is a clear path to continued access to the data, and the data can be made available to end users for free.
  • Point predictions without uncertainty. We prefer tools that present results as estimates, ranges, or sensitivity analyses rather than single-point forecasts that claim precision.
  • Unverified AI-generated code. AI-assisted development is welcome, but Submissions must be tested, explainable, and secure. Code that the Team cannot explain, that lacks tests of its core logic, or that ships with unaddressed security findings will not score well.

Benchmark Policy Scenarios

To enable meaningful comparison and consistent judging across Submissions, all tools must estimate a baseline scenario, which models housing production under the current regulatory and economic environment without policy changes. Teams should assume a five year time horizon. Teams may choose the housing outcome they model (e.g., permits, starts, completions, or housing stock growth), but must clearly define and justify that choice.

Teams must also estimate the change in housing production or aggregate housing stock under at least three of the five scenarios listed below. For each scenario addressed, Teams should demonstrate:

  • How the scenario changes the tool’s outputs relative to the baseline;
  • What drives the differences (assumptions, data, model mechanics);
  • How sensitive the results are to key inputs and assumptions; and
  • What the model cannot tell policymakers about the scenario.

Baseline scenario

Housing production under the current regulatory and economic environment, with no policy changes.

The five scenarios

1. Accessory Dwelling Units (ADUs)

Model the effect of allowing one ADU on parcels with existing single-family buildings. Assume detached and semi-detached ADUs of at least 600 square feet are allowed, and no additional parking is required.

2. Missing Middle/Gentle Density

Model the effect of allowing 2–6 dwelling units on parcels currently restricted to single-family use, with building heights up to three stories. This scenario reflects the “gentle density” reforms (e.g., townhouses, 2–6-unit multifamily buildings) advancing across states and cities nationwide. Teams may assume that setback and lot coverage requirements do not constrain the development of 2–6 units on a parcel currently restricted to single-family use.

3. Transit-Oriented Density

Model the effect of increasing allowable residential density near major transit stops through higher building heights and/or increased floor area ratios. Teams should define major transit stops and justify their transit proximity threshold. Teams may assume that parking requirements are reduced or eliminated.

4. Parking Requirement Reduction

Model the effects of reducing or eliminating minimum parking requirements for new residential development. Specify assumptions about developer-provided parking in the absence of mandates.

5. Development Fee Reduction

Model the effect of reducing total development fees (planning, environmental review, building, and impact fees) by 50% across all major fee categories. For comparability across submissions, Teams should assume development fees total $50,000 per unit across all major fee categories.

Teams are encouraged to develop additional state-specific scenarios with mentor guidance, particularly those reflecting state-specific policy debates in the chosen jurisdiction.

Technical, Reproducibility, and Security Guidance

Starter template and reference. To put every Team on equal footing, the Challenge platform provides a ready-to-use template repository and a reference example (see Starter Materials and Resources below). The template pre-configures the required workflows and documentation, so you inherit the setup rather than building it. Using the template is strongly recommended but not required, provided your Submission satisfies the required Technical, Reproducibility, and Security Requirements as set forth in the Official Rules.

Languages. Any memory-safe language is acceptable (e.g., Python, Java, Go, and others). Where a required tool does not support your language (e.g., CodeQL), rely on the remaining tooling and document the gap.

Data and Resources

Provided data. Teams should use available real-world datasets, which may include satellite data, zoning and land-use maps, parcel characteristics, and building permit histories.

Recommended public data sources. Teams are encouraged to use nationally available datasets, including but not limited to:

  • U.S. Census Bureau/American Community Survey (ACS)
  • Bureau of Labor Statistics (construction employment, costs)
  • Land cover and land use datasets, including the USGS National Land Cover Database (NLCD) and the NASA Land Use/Land Cover database
  • Building footprint datasets, such as Microsoft’s Building Footprints and Google’s Open Buildings
  • Census Building Permits Survey
  • Federal Housing Finance Agency House Price Index (FHFA HPI)
  • USPS vacancy data
  • National Transit Database
  • State, local, tribal, and territory open data portals (zoning, parcel, assessment data)

Team-sourced data. Teams may augment the provided datasets with additional publicly available data.

Provenance. Document every dataset’s source, version, access date, license, cleaning steps, and known limitations in your Model & Assumptions Overview and the data-provenance template.

Sensitive and restricted data. Do not use sensitive, confidential, restricted-use, or personally identifiable data. If you use de-identified or aggregated data, describe the method and any re-identification or small-cell risk.

Embedding data. Teams must not include large source datasets directly in their Submission. Instead, Submissions must include an automated data acquisition process that retrieves all required datasets from their authoritative source and reproduces the analysis environment with minimal manual intervention.

You are solely responsible for complying with any and all license terms related to such material.

Starter Materials and Resources

The Challenge Platform provides:

  • Template repository — a pre-configured GitHub template with the required workflows (CI, CodeQL, Dependabot), documentation templates, and a submission checklist.
  • Reference example — a small, transparent example tool that passes every gate, illustrating the expected structure and rigor (not a methodology or language endorsement).
  • Document templates — Model & Assumptions Overview, Policy Scenario Comparison, Data Provenance, and Disclosures.
  • Repository setup guide — steps to enable secret scanning, push protection, CodeQL, Dependabot, and SBOM export.
  • Curated dataset list — starting-point datasets with sources and access notes.

Publication and Post-Challenge

Public gallery. Submissions may be featured in a public project gallery after winners are announced. Because all submissions are open source, their repositories remain publicly available.

Post-challenge report. Sponsor or Administrator may publish a report summarizing outcomes and lessons learned.

Continued availability. Winning teams are encouraged to keep their repositories available so the tools can inform ongoing policy work.

Frequently Asked Questions

Can I use any programming language? Yes. Any language is acceptable, provided your Submission passes the eligibility gates.

Do I have to use the starter template? No, but it is strongly recommended—it pre-configures the required security and reproducibility setup.

Can I use AI tools to write code? Yes. You must be able to explain, test, and secure the result, and disclose material use.

Which scenarios must I model? The baseline plus at least three of the five benchmark scenarios.

Does open-sourcing my code mean I give up ownership? No. An open-source license grants use rights; it does not transfer ownership.

Can I use paid or proprietary data or APIs? Yes, if there is a clear path to continued free access for end users. In addition, you must disclose all such dependencies.

Glossary

  • ADU (Accessory Dwelling Unit): a secondary dwelling on a lot with a primary home.
  • Missing middle: housing types between single-family homes and large apartment buildings (duplexes, townhouses, small multiplexes).
  • Floor area ratio (FAR): the ratio of a building’s floor area to the size of its lot.
  • Pro forma: a financial model estimating a development project’s costs, revenues, and feasibility.
  • Baseline: modeled housing production under current rules, with no policy change.
  • Eligibility gate: a pass/fail requirement a submission must meet before it is scored.
  • CodeQL: GitHub’s code-scanning tool that detects security vulnerabilities in source code.
  • Dependabot: GitHub’s tool that flags and updates dependencies with known vulnerabilities.
  • Secret scanning: GitHub’s detection of committed credentials.
  • SBOM (Software Bill of Materials): an inventory of a project’s software dependencies.
  • OSI-compliant license: a license approved by the Open Source Initiative.

Submission Checklist

This Checklist is advisory and is not a guarantee that a Submission will be deemed complete. Refer to the Official Rules at Official Rules for all Submission requirements.

Eligibility gates

  • ☐ Reproducible setup from documented commands
  • ☐ Automated tests covering core logic, with a documented run command
  • ☐ No committed secrets; secret scanning enabled, zero unresolved alerts
  • ☐ Dependabot alerts and CodeQL enabled; no unaddressed critical/high findings
  • ☐ Approved OSI-compliant license (Apache-2.0 default; MIT/BSD-2/BSD-3 allowed)

Deliverables

  • ☐ Interactive tool (functional)
  • ☐ Slide deck (≤ 30 slides, PDF)
  • ☐ Demo video (5 minutes, viewable without an account)
  • ☐ Model & Assumptions Overview (≤ 25 pages)
  • ☐ Source code (public GitHub repo)
  • ☐ Software Bill of Materials (SBOM)
  • ☐ Policy scenario comparison (≤ 4 pages)

Scope and disclosures

  • ☐ Baseline + at least three benchmark scenarios modeled at the state level
  • ☐ Data provenance documented
  • ☐ AI tools, models, APIs, and dependencies disclosed