For a few values, we can write them directly in code.
mass = 5.97f;
distance = 150f;
velocity = 29.8f;
But what if we have:
10 objects? → 1,000 objects? → 1,000,000 objects?
For large datasets, writing every value directly in code is no longer practical.
Instead, we store (or obtain) the data in an external file:
DATA FILE → Program → Simulation
Example:
name,mass,distance,velocity
Earth,5.97,150,29.8
Mars,0.642,228,24.1
Jupiter,1898,778,13.1
Scalability
Use the same code with 10, 1,000, or 100,000 data records.
Separation of Data and Code
Change the data without rewriting the simulation code.
Reusability
The same simulation can run with different datasets.
Maintainability
Data can be updated, generated, or replaced independently.
Keyword: Table
name,mass,distance
Earth,5.97,150
Mars,0.642,228
Good for:
Many records with the same structure
Numerical and tabular data
Excel / spreadsheets
Simple large datasets
- Original document: rfc-editor.org/rfc/rfc4180.html
Keyword: Structure
{
"bodies": [
{ "name": "Sun", "type": "Star", "mass": 199000000, "orbitAU": 0, "distance": 0, "initial_velocity": 0 },
{ "name": "Mercury", "type": "Planet", "mass": 33, "orbitAU": 0.387, "distance": 57894375961, "initial_velocity": 47360 },
{ "name": "Venus", "type": "Planet", "mass": 487, "orbitAU": 0.723, "distance": 108159260516, "initial_velocity": 35020 },
{ "name": "Earth", "type": "Planet", "mass": 597, "orbitAU": 1, "distance": 149597870700, "initial_velocity": 29780 },
{ "name": "Mars", "type": "Planet", "mass": 64.2, "orbitAU": 1.524, "distance": 227987154947, "initial_velocity": 24070 },
{ "name": "Jupiter", "type": "Planet", "mass": 190000, "orbitAU": 5.203, "distance": 778357721252, "initial_velocity": 13070 },
{ "name": "Saturn", "type": "Planet", "mass": 56800, "orbitAU": 9.537, "distance": 1426714892866, "initial_velocity": 9680 },
{ "name": "Uranus", "type": "Planet", "mass": 8680, "orbitAU": 19.191, "distance": 2870932736604, "initial_velocity": 6800 },
{ "name": "Neptune", "type": "Planet", "mass": 10200, "orbitAU": 30.069, "distance": 4498258374078, "initial_velocity": 5430 },
{ "name": "Pluto", "type": "Dwarf", "mass": 1.3, "orbitAU": 39.482, "distance": 5906423130977, "initial_velocity": 4740 }
]
}
Good for:
Structured data
Objects and arrays
Nested relationships
Configuration files and web data
JSON is a text-based format for representing structured data and can contain objects and arrays.
Official documentation: Working with JSON - Learn web development | MDN
CSV
Many things with the same variables.
JSON
Things with more complex structures and relationships.
They are:
Simple · Human-readable · Easy to create · Easy to import
and provide a good introduction to:
Data → Code → Simulation
CSV and JSON are both text-based formats.
They are excellent places to start, but:
If we have millions or billions of numbers, we may eventually stop using both.
For very large or complex datasets, we may use:
Binary formats · Databases · HDF5 · Specialized data systems
HDF5, for example, combines a binary file format with a data model for organizing datasets and groups.
example) Ocean data over the globe
About HDF5: HDF5: Introduction to HDF5
CSV: Think Spreadsheet
JSON: Think Objects / Tree
HDF4: Think Large Scientific Data Container
Tabular Data
Thousands of discovered planets with properties such as orbital period, radius, and mass.
https://exoplanetarchive.ipac.caltech.edu/
Structured / Nested Data
Real-time earthquake data including location, depth, magnitude, and time in GeoJSON format.
https://earthquake.usgs.gov/earthquakes/feed/
Large Scientific Data
Satellite observations containing large multidimensional arrays, spectral bands, geolocation, and metadata.
https://modis.gsfc.nasa.gov/data/