CodeSage AI
AI-Powered Technical-Debt Analytics Dashboard
Abstract
CodeSage AI is an AI-powered technical debt analytics platform made for small-scale agile software development teams. It analyzes the source code and development history of the codebases using static analysis, behavioral analysis metrics, and machine learning to identify, track, and prioritize technical debt based on its likelihood of contributing to future bugs. This helps the small-scale agile software deveelopemt teams to focus on the technical debt that matters most instead of every detected techincal-debt. CodeSage helps them focus on the issues that matter most, allowing them to reduce risk while continuing to deliver new features efficiently.
The problem
Small-scale agile software teams must continuously balance code quality against the pressure to deliver new features as fixing a technical debt usually does not create a visible change for the client. Technical debt accumulates as software evolves and fixing some of the technical debt can cost more time as well as effort. However, treating every detected technical debt as equally important creates large, noisy backlogs and can consume development time on low-impact problems.
Therefore, teams need more than existing traditional static-analysis results and they need a way to understand which of the detected technical-debt issues are most important, which files are more likely to experience future defects, and where limited refactoring effort should be focused. CodeSage AI addresses this decision-making gap by helping teams determine what should be fixed first, rather than simply identifying everything that could be fixed. The system is specifically intended to let developers balance debt fixing with feature development and release commitments.
The solution
CodeSage AI is a web-based, multi-tenant technical-debt analytics platform that connects to a GitHub repository and analyzes a selected revision using static analysis, repository-history analysis, deterministic rules, and two specialized machine-learning models. Repository scans run asynchronously through Celery workers and Redis so that lengthy analysis does not block the user interface.
For each scan, CodeSage extracts static software metrics using CK, process metrics such as churn, author count, file age, and recency using PyDriller, and source-code comments using Tree-sitter. A deterministic rule engine identifies code-design and security findings. The first ML model, ML-1, analyzes source-code comments to detect Self-Admitted Technical Debt and classify it into code-design, requirement, documentation, or test debt. The second model, ML-2, combines static and process metrics to estimate each file's bug-proneness. Unlike ML-1, ML-2 does not create findings; instead, its risk prediction influences how existing findings are prioritized.
The system then combines each finding's severity, debt-category importance, source confidence, recent code churn, and predicted file risk to calculate a priority score. Developers can adjust category weights and the balance between rule-based and model-based evidence through scoring profiles, and CodeSage recalculates priorities from the stored analysis without requiring another repository scan.
These results are presented through an interactive dashboard containing an overall health score and grade, category breakdown, historical health trends, hotspot file tree, detailed evidence, and a ranked Refactor-First list showing the technical-debt findings that deserve attention first.
Screenshots
Related projects
Automated Code Review and Technical Debt Tracking Dashboard
Automates Code Review and Tracks Technical Debt From Pull Requests