|
Title: Performance Engineer
Duration: 12 Months (may convert to FTE)
Location: Southlake, TX or Ann Arbor, MI, hybrid onsite 4 days Job Description:
Client is seeking a Senior Performance Engineer to join the thinkorswim Performance Engineering team. This individual-contributor role independently plans and delivers performance and resiliency initiatives for a Java-based trading platform, with a focus on latency, scalability, reliability, and resiliency.
The engineer builds workloads and tools, finds code and system bottlenecks, and recommends improvements across client, server, gateway, infrastructure, and cloud environments.
This role requires deep Core Java knowledge and strong diagnostic skills. The engineer reads source code, analyzes JVM and system behavior, and evaluates production traffic patterns. The engineer translates those findings into recommendations for code, configuration, capacity, and system design.
Key Responsibilities
* Responsible for independently planning and delivering performance, scalability, resiliency, and capacity validation for assigned systems.
* Learn architecture, business flows, dependencies, workloads, and production behavior.
* Define requirements, workloads, acceptance criteria, and reports with partner teams.
* Build performance test scripts, workload generators, mock services, frameworks, and automation.
* Run load, stress, endurance, scalability, failover, recovery, degradation, and resiliency tests.
* Analyze source code, JVM behavior, system resources, databases, networks, logs, metrics, and end-to-end latency.
* Recommend and validate code, configuration, infrastructure, database, capacity, design, and resiliency improvements(CircuitBreakers, Shapers, LoadBalancers & Failover).
* Use production behavior as the baseline for workload design, capacity planning, risk analysis, and validation.
* Maintain dashboards, reports, diagnostic tools, runbooks, frameworks, and workload libraries.
* Evaluate new testing, profiling, observability, and resiliency tools.
* Use approved automation and AI tools to improve scripting, analysis, documentation, and diagnostics.
* Communicate results, risks, recommendations, and open issues clearly.
Required Qualifications
* Bachelor's degree in computer science or engineering.
* 5+ years in performance, scalability, software engineering, or production performance analysis.
* Strong experience with performance, scalability, failover, recovery, and resiliency testing.
* At least 2 years of hands-on Java development experience.
* Deep Core Java knowledge, including JVM internals, concurrency, multithreading, memory, garbage collection, and latency. Able to analyze source code and recommend code-level improvements.
* Ability to read source code and diagnose bottlenecks with dumps, logs, profilers, metrics, and observability tools.
* Experience building workload generators, utilities, mock services, or automation with Java, Scala, Python, shell, JMeter, Gatling, or similar tools.
* Strong knowledge of Linux, databases, application servers, caching, connection pools, distributed systems, and networks.
* Ability to create workloads, baselines, thresholds, capacity measures, acceptance criteria, SLOs, SLAs, and KPIs.
* Ability to provide evidence-based code, configuration, capacity, reliability, and design recommendations.
* Good experience in using Splunk, Grafana (or other monitoring tools)
* Experience in testing resiliency solutions by updating CircuitBreakers, Shapers, LoadBalancers & Failover configs.
* Experience with troubleshooting/diagnosing JVM issues (e.g. thread dumps, garbage collection
and memory management)
* Understanding of performance best practices, performance key metrics, and statistics
* Ability to manage priorities and work with distributed teams.
Preferred Qualifications
? Real-time financial systems(trading) experience is a big plus
? Google Cloud, Kubernetes, or Docker is preferred
? AI-assisted analysis, LLM tools, or intelligent automation
? Usage experience of AI tools with Visual Studio/ Idea IntelliJ is a plus
|