We collect the data.
Then we do the
research.
Data Tune supplies the raw material behind other people's research — and takes on a small, deliberately limited set of research subjects ourselves. Web scraping, data mining, natural language processing and machine learning, run by a team that publishes its own work openly so you can check it before you commission anything.
How the work flows
Two ways to work with us
Most clients need data, not a researcher. A few need both. We are set up for either, and we will tell you honestly which one your question needs.
1 — We supply the data, you do the research
This is the bulk of our work. You define the question; we design the collection method, build it, run it and hand over a clean, documented dataset your own analysts or supervisors can work from.
- Data collection, online and physical
- Custom web scraping and API extraction
- Cleaning, de-duplication, labelling and validation
- Automated pipelines that keep refreshing
- Dashboards so the data stays visible
2 — We take the research outsourcing
We accept full research briefs, but only inside a short list of subjects where we have real depth. Outside that list we would be guessing, and we would rather say so than sell you a report we cannot stand behind.
- Question framing and method design
- Collection, mining and statistical analysis
- A written report with every figure reproducible
- The dataset and the scripts handed over with it
Four things we are genuinely good at
These are the capabilities every project we take on rests on. If your brief does not need at least one of them, we are probably not the right supplier.
Data mining
Classification, clustering, prediction and pattern discovery over unstructured sources — turning noise into something a decision can rest on.
Web scraping
Custom scrapers and directory extraction built to survive layout changes, with rate limiting, completeness checks and lawful, documented scope.
NLP
Natural language processing on real-world text, including code-mixed Sinhala and English — sentiment, topic, intent, entity extraction and classification.
Machine learning
Model-ready labelled datasets, annotation to a written standard, and applied ML support where a model is the right answer rather than the fashionable one.
Pipeline or waterfall
Two ways the data reaches you. Most engagements are one or the other; a few start as waterfall and become a pipeline once the shape of the data is settled.
Pipeline — continuous
A feed that keeps running. Sources are pulled on a schedule, cleaned, validated and written into your database or dashboard automatically. You watch the numbers move rather than waiting for a delivery date.
- Scheduled pulls with rate limiting and back-off
- Completeness receipts on every run
- Alerting when a source breaks or drifts
- Best for monitoring, tracking and live dashboards
Waterfall — fixed scope
A defined question, a fixed scope, staged milestones and one delivery at the end. Scope is agreed in writing before anything starts, and you get scheduled progress updates with sample batches along the way.
- Written scope, volume and price agreed up front
- Staged progress updates with sample data
- Single validated handover of everything produced
- Best for one-off studies, audits and thesis support
The subjects we take on
Deliberately short. These are the areas where we have run the work before, hit the problems, and know what the data does when it misbehaves.
Social experiments
Behavioural and opinion studies built on observed online activity or structured field collection — sentiment, response to messaging, adoption patterns, community reaction.
Image analysis
Image processing and computer-vision studies: classification, detection, shelf and condition auditing, labelled image sets built and validated to a written standard.
IoT projects
Sensor deployment and the research built on top of it — environmental monitoring, device telemetry, calibration studies and the pipelines that keep the readings honest.
Telecommunications
Network, usage and service-quality data. Collection, structuring and analysis of telecom datasets, an area our founder has worked in directly for years.
What we will take on
- Research inside the four subjects above, end to end
- Data collection and mining for any subject or industry
- Dataset building, labelling and validation for ML
- Academic and thesis data support in any faculty
- Pipelines, dashboards and tracking tools
What we will not
- Full research in subjects outside our four — we will supply the data instead
- Anything requiring collection that is not lawful and consent based
- Personal data harvesting or scraping behind a login
- Work where the conclusion is decided before the data is collected
- Volumes we cannot deliver inside the agreed window
Read our research before you commission any
We publish our studies in full — the report, the underlying dataset and the scripts that regenerate every figure in it. Free, no signup, no email form. Check the method yourself, then decide whether to hire us.
How an engagement runs
Four steps, and you can stop after the first two without owing us anything.
Consultation
Tell us the question, the subject and the deadline. We say plainly whether it is a data job or a research job, and whether we are the right people for it.
Scope & quote
A written scope with the collection method, the volume, the delivery model and a firm price. Free, and no obligation to continue.
Collection & mining
We build and run it, with scheduled progress updates and sample batches so nothing is a surprise at the end.
Delivery & support
Dataset, documentation and, where the brief includes it, the report and the scripts. Support afterwards so the data actually gets used.
Data integrity, ethics and compliance
All data is gathered lawfully, transparently and on a consent basis, with scope agreed before work begins. Our handling aligns with the EU General Data Protection Regulation (GDPR), applicable United States frameworks such as the CCPA, and Sri Lanka's Personal Data Protection Act No. 9 of 2022 — data minimisation, secure storage, anonymisation of personal identifiers and strict client confidentiality.
Tell us what you need data on
If your subject is on our list we will quote the whole study. If it is not, we will quote the data and point you to the right analyst. Either way you get a straight answer and a free quote.
Fourteen years in the IT industry on the service and engineering side, and five years working specifically as a data solutions consultant across image processing, telecommunications data, IoT development and data-driven research.
Every brief is read personally. If we are not the right fit for your subject, you will be told that in the first reply rather than three weeks in.