Builds robust, fault-tolerant data pipelines that collect, assemble and potentially transform and aggregate unorganized data distributed into databases or data sources such as operational data stores, data integration hubs and data lakes or estuaries.
Compiles and installs database systems, writes queries, scales to multiple machines and puts disaster recovery systems into place.
Builds groundwork for data consumers (software or human) to easily retrieve needed data for evaluations and experiments.
Builds operational data use cases such as moving large volumes of data across applications via operational data stores, data hubs and data lakes, and builds private/segregated data pipelines between specific applications.
Designs and implements scalable, reliable distributed data processing frameworks and analytical infrastructure using multiple technologies, including data sets or data warehouses, data virtualization and services and repositories of semi-structured data sets.
Defines, designs, and implements data management, storage, backup and recovery solutions that ensure high performance of the organization's enterprise data. Designs automated software deployment functionality that allows efficient management of applications across distributed platforms.
Understands structural requirements and defines standards for how data will be stored, consumed, integrated and managed. Monitors structural performance and utilization, identifies problems and implements solutions.
Lead the creation of standards, best practices and new processes for operational integration of new technology solutions.
Ensures environments are compliant with defined standards and operational procedures. Implements measures to ensure data accuracy and accessibility, constantly monitoring and refining the performance of data management systems.
Completes problem tickets including bug fixes, design modification and enhancement based on customer requirements.
EDUCATION:
Bachelor's Degree or equivalent combination of education and experience required.
REQUIRED SKILLS AND ABILITIES:
Hadoop-based technologies including MapReduce, Spark, Hive, Presto and Pig. SQL based technologies such as Oracle, PostgreSQL and MySQL. Python or equivalent high-level language.
Understanding of NoSQL technologies like Cassandra and MongoDB.
Understanding of data warehousing solutions.
Experience building distributed and cloud-based data pipelines.
Understanding of industry standard software APIs.
Strong customer service skills. Organization and time management skills. Business consulting abilities.
Commitment to system functionality and user satisfaction in rapidly growing and changing environments.
Demonstrated initiative in resolving problems, balancing conflicting requirements in partnership with others.
Ability to function in a fast-paced high energy environment. Familiarity with Agile and Scrum methodologies.
PHYSICAL DEMANDS:
Extensive sitting, phone and computer use. Some travel may be required.
WORK ENVIRONMENT:
General office environment. Normal office noise level, with occasional moderate noise.