Data radar: unclassified entries
Momentum based on available 7-day snapshots for Data: measured with two snapshots, estimated otherwise. Each value shows its own source window. This page reads stored snapshots only.
Radar filters
Tool ranking
Ranked by measured growth, normalized across sources. Raw values keep their source-specific window, shown under each tool; a change requires at least two snapshots from the last 7 days.
- 1Active
End-to-end streaming data platform on AWS: Kafka → Iceberg lake → Kimball model → Athena, orchestrated with Airflow 3. Terraform, Glue PySpark, dbt.
githubmeasured growthOpen source ↗
106GitHub stars+51 (+92.7 %) - 2Active
semantica
OtherGraph-Native Infrastructure for Context and Accountable AI Systems
githubmeasured growthOpen source ↗
11 532GitHub stars+698 (+6.4 %) - 3Active
dify
OtherBuild Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
githubmeasured growthOpen source ↗
154 043GitHub stars+524 (+0.34 %) - 4Active
Summer 2027 software engineering, data science, AI, quant, product management, and hardware internship postings. Updated daily by Simplify and Pitt CSC.
githubmeasured growthOpen source ↗
47 002GitHub stars+196 (+0.42 %) - 5Active
data-eng-bench
OtherData-engineering benchmark for coding agents (DuckDB + Snowflake dbt tasks).
githubmeasured growthOpen source ↗
64GitHub stars+11 (+20.8 %) - 6Active
openobserve
OtherOpen source observability platform for logs, metrics, traces, frontend monitoring, pipelines and LLM observability. A sophisticated, simple and highly performant alternative to Datadog, Splunk, and Elasticsearch with 140x lower storage…
githubmeasured growthOpen source ↗
21 594GitHub stars+119 (+0.55 %) - 7Active
superset
OtherApache Superset is a Data Visualization and Data Exploration Platform
githubmeasured growthOpen source ↗
74 566GitHub stars+103 (+0.14 %) - 8Active
Tech internships & CS jobs tracker - 180 open software engineering internships from 4,300 employer job boards, auto-updated every 30 min with visa-sponsorship and H-1B sponsor data.
githubmeasured growthOpen source ↗
698GitHub stars+26 (+3.9 %) - 9Active
harness
OtherHarness Open Source is an end-to-end developer platform with Source Control Management, CI/CD Pipelines, Hosted Developer Environments, and Artifact Registries.
githubmeasured growthOpen source ↗
38 196GitHub stars+72 (+0.19 %) - 10Active
dbt-core
Otherdbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.
githubmeasured growthOpen source ↗
13 750GitHub stars+60 (+0.44 %) - 11Active
lucebox
OtherLLM speculative inference server for heterogeneous hardware & consumer GPUs
githubmeasured growthOpen source ↗
2 825GitHub stars+37 (+1.3 %) - 12Active
haystack
OtherOpen-source AI orchestration framework for building context-engineered, production-ready LLM applications. Design modular pipelines and agent workflows with explicit control over retrieval, routing, memory, and generation. Built for…
githubmeasured growthOpen source ↗
26 377GitHub stars+60 (+0.23 %) - 13Active
2026–2027 tech internships in the US, updated daily. Software engineering, data science, security, product, and design intern roles.
githubmeasured growthOpen source ↗
93GitHub stars+8 (+9.4 %) - 14Active
airflow
OtherApache Airflow - A platform to programmatically author, schedule, and monitor workflows
githubmeasured growthOpen source ↗
46 672GitHub stars+64 (+0.14 %) - 152 074GitHub stars+29 (+1.4 %)
- 16Active
argo-cd
OtherDeclarative Continuous Deployment for Kubernetes
githubmeasured growthOpen source ↗
24 046GitHub stars+54 (+0.23 %) - 17Active
duckle
OtherOpen-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
githubmeasured growthOpen source ↗
1 254GitHub stars+24 (+2.0 %) - 18Active
spark
OtherApache Spark - A unified analytics engine for large-scale data processing
githubmeasured growthOpen source ↗
43 931GitHub stars+55 (+0.13 %) - 19Active
spark-vllm-docker
OtherDocker configuration for running VLLM on dual DGX Sparks
githubmeasured growthOpen source ↗
2 205GitHub stars+27 (+1.2 %) - 20Active
kestra
OtherEvent Driven Orchestration & Scheduling Platform for Mission Critical Applications
githubmeasured growthOpen source ↗
27 967GitHub stars+45 (+0.16 %) - 2122 497GitHub stars+41 (+0.18 %)
- 226 110GitHub stars+29 (+0.48 %)
- 23Active
dlt
Otherdata load tool (dlt) is an open source Python library that makes data loading easy 🛠️
githubmeasured growthOpen source ↗
5 804GitHub stars+26 (+0.45 %) - 24Active
great_expectations
OtherAlways know what to expect from your data.
githubmeasured growthOpen source ↗
11 760GitHub stars+25 (+0.21 %)
Learning resources
Ranked by measured growth, normalized across sources. These resources remain available separately and do not take part in the main tool ranking.
- 1ActiveOpen source ↗
Data Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. Join the course here 👇🏼
githubResourcemeasured growth45 131GitHub stars+171 (+0.38 %)