DBT-BUILDER · Group 3 · JSS AHER
Echo reports go in as PDFs.Data comes out as a table.
EchoMiner converts semi-structured echocardiography PDF reports into analysis-ready structured data. Developed under the DBT-BUILDER project at JSS Academy of Higher Education and Research, Mysore.
- PDFs per batch
- 5
- fields per report
- 47
- files kept after download
- 0
PDFs per batch
fields per report
files kept after download
Doppler study · spectral envelope
Illustrative values. Not from a patient record.
About
What EchoMiner does
Echocardiography reports are written for clinicians to read, not for software to analyse. EchoMiner reads them the way a research assistant would, and returns a table.
A departmental echo archive holds thousands of PDF reports. Each one contains the same measurements in roughly the same places, but as prose and layout rather than as data. Extracting them by hand is the step that stops most retrospective studies before they start.
EchoMiner applies a validated rule-based pipeline to those reports and returns one row per study, with the measurement fields, the narrative findings and the impression lines in separate columns — plus a quality summary of what it read from each file.
Rationale
Why structured echocardiography reports matter
Free text cannot be counted
A cohort described in paragraphs cannot be summarised, stratified or modelled without first being turned into variables.
Manual entry does not scale
Transcribing an archive by hand is slow, and every pass introduces variation that is invisible in the final dataset.
Reproducibility needs provenance
A deterministic pipeline with a recorded version lets a reviewer regenerate the same dataset from the same reports.
Features
What you get
Batch extraction
Upload up to 5 echocardiography PDFs at once and receive one consolidated workbook.
Rule-based and reproducible
Deterministic regular-expression extraction. The same input always produces the same output, and every workbook records the engine version that produced it.
Quality reporting
Every export states pages read, reports detected in the source, rows returned, and any parse warnings, per file.
Analysis-ready output
Measurements arrive as numeric cells with a data dictionary sheet, ready for statistical software.
Nothing retained
Uploaded PDFs and generated files are deleted as soon as your download completes.
Citable
Archived on Zenodo with a DOI. Citation formats are included in every workbook.
Workflow
Five steps, one page
Registration and the tool live on this page. Nothing redirects you elsewhere.
- 01
Register once
Verify your email with a one-time code. Subsequent visits from the same device skip it.
- 02
Upload reports
Drop up to 5 PDF reports. Files are checked before anything is read.
- 03
Extract
The validated pipeline reads each report and maps it to the structured field set.
- 04
Review
Preview the extracted table and the per-file quality summary in the browser.
- 05
Download and clear
Take the branded workbook. Your files are removed from the server immediately.
Access
Launch EchoMiner
Register once, verify your email, and the tool opens right here on this page.
Already registered?
Funding
JSSAHER DBT BUILDER Project
JSS Academy of Higher Education & Research (JSSAHER) has been selected by the Department of Biotechnology (DBT) to implement the prestigious DBT BUILDER (Boost to University Interdisciplinary Life Science Departments for Education and Research) program.
Backed by a ₹5 crore grant over five years, this initiative promotes interdepartmental collaboration to nurture postgraduate talent and build a globally competitive bio-economy.
The project targets the prevention and management of cardiovascular diseases across three emerging research domains:
- 1Novel Biomarker and TherapeuticsMetabolic disorders and cardiopulmonary disease.
- 2NanotheranosticsAdvancing CVD disease management.
- 3Spatial Health Informatics and ManagementThe domain EchoMiner is built in.
Spatial Health Informatics & Management (Group 3)
Group 3 leads the spatial health informatics domain and is the driving force behind the development of AI_EchoMiner — a scalable, Python-based data extraction framework that utilizes regular expressions and pandas to process complex healthcare data efficiently.


People
Project leadership and development team
Project leadership
Dr. Rajesh Kumar Thimmulappa
Principal Investigator | Professor, Dept. of Biochemistry, JSS Medical College
kumar_rt@yahoo.comDr. Madhu B
Co-Principal Investigator | Professor & Head, Dept. of Community Medicine, JSS Medical College
madhub@jssuni.edu.in
Group 3 research & development team
Dr. Manjunatha M C
Assistant Professor, Dept. of Community Medicine, JSS Medical College, Mysuru
mcmanju1@gmail.comSuraj B M
Senior Research Fellow, Dept. of Community Medicine, JSS Medical College, Mysuru
surajbm@jssuni.edu.in
Impact
Research impact
EchoMiner exists to make retrospective echocardiography research feasible at archive scale within the DBT-BUILDER cardiovascular programme.
Usage and output metrics are published here once verified by the project team. Figures are drawn from platform records rather than estimates, and this section stays empty until they are confirmed.
Output
Publications
Publications arising from EchoMiner will be listed here as they appear. The software itself is archived and citable:
B Manjunath, S. (2026). EchoMiner: source code for rule-based NLP extraction from echocardiography PDF reports [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.21281483
https://doi.org/10.5281/zenodo.21281483Updates
News and updates
No updates have been posted yet. Project announcements, releases and new report-layout support will appear here.
Attribution
How to cite EchoMiner
Acknowledge EchoMiner in any publication, thesis, conference paper, report or scientific communication that uses data generated through this tool.
B Manjunath, S. (2026). EchoMiner: source code for rule-based NLP extraction from echocardiography PDF reports [Computer software]. Zenodo. https://doi.org/10.5281/zenodo.21281483
Questions
Frequently asked questions
What file formats can I upload?
PDF only, up to 5 files per submission and 50 MB in total. Up to 10,000 pages per submission — a quarterly archive usually fits in one. The PDF must contain a text layer — scanned images without OCR cannot be read.
Do I need to verify my email every time?
No. A one-time code is required at first registration. After that, returning from the same browser restores your access silently. A new device or browser asks for one code again.
What happens to the reports I upload?
They are held only while your job runs and are deleted the moment your download completes. Jobs that are abandoned are swept automatically. Usage records such as file counts and timestamps are retained; report content is not.
Is the extraction accurate?
The pipeline is rule-based and deterministic, and it is validated against the report layouts it was built for. Every workbook includes a Quality sheet showing what was read from each file so you can verify the output rather than assume it.
How should I cite EchoMiner?
Use the citation in the How to Cite section, also included in every exported workbook in APA, Vancouver, IEEE, BibTeX and RIS.
Who can use EchoMiner?
Registration is open to researchers. Users are responsible for holding the appropriate ethical approvals for any data they process through the tool.
Contact
Get in touch
Dr. Madhu B
Professor & Head
Co-Principal Investigator
Department of Community Medicine
JSS Medical College
JSS Academy of Higher Education and Research
Mysore, India
madhub@jssuni.edu.inData privacy
The information collected through this portal is used solely for providing access to the EchoMiner research platform and maintaining institutional usage records. User information will be securely stored in accordance with applicable Government of India data protection guidelines and institutional policies. The information will not be shared with any third party except where required by law or institutional policy.