VedexVendorsDataMynd

DataMynd

View in Graph View in Semantic Map

io This native application is designed for users who want a quick way to generate synthetic data that is based on one or multiple tables in a real production schema. The application works by training one of several different ML models on the real data before using the trained model to generate synthetic data. The application performs all logic within the native application on the consumer's Snowflake account.

No external API calls are made at any point either during installation or use. No third-party services are used at any point by the application. No Data or metadata are shared outside of the consumer's account.

Privileges The application requires the following grants to access the targeted data and train the model. GRANT USAGE ON DATABASE GRANT USAGE ON SCHEMA GRANT REFERENCES, SELECT ON TABLE Workflow [Start Page] Create a new project or resume an existing project. The first page of the app contains a list of all open projects and high-level stats on each, like the schema/tables/chosen learning model and whether the model has been trained successfully.

Clicking on a project will give the user options to resume or delete the project. Note: deleting the project will also drop any synthetic data generated using that project. Please copy the data before doing this.

[Select Page] The user is prompted to select the real data for the project. Data selection is done by using the drop-down lists shown on this page. Note: the user is prompted when creating the project to grant access to the application for the desired tables in the previous step.

Failure to do so will result in a warning and the user will not be able to proceed from this step. [Select Page] On the same page the user selects the real data, the learning model must also be selected. Different models vary in accuracy and speed for different use cases and data size/complexity/shapes.

GaussianCopula uses purely statistical methods when learning and performs the fastest with lowest accuracy. CTGAN and TVAE use neural-networks and deep learning when profiling the real data and can result in highly accurate synthetic data. CopulaGAN uses a combination of statistical and deep learning methods.

The neural network-based models take a longer time to train but will likely produce more accurate, realistic results. [Configure Page] The user must select the table, then configure several options for each column in the dataset. One primary key should be selected (using the P-Key checkbox) if available for each table.

"ID" must be selected as the type for the primary key. The type attribute should match that of the original data. Categorical should be selected as the type if the field contains discrete values (text values like 'red' or 'green' or distinct numerical values like radio station frequency that do not make sense to plot as a numerical distribution).

Otherwise, one of the other types should be selected corresponding to the datatype of the original. [Configure Page] The user has the option to select 'anonymize' for any categorical-type columns. This will cause the model training to skip over that field and instead it will create artificial values using the Faker library.

If anonymize is selected, the user must also select the type of output the user expects to replace the field values for that column. g. selecting 'name' for the type (person category) will randomly generate full names to populate the field values.

Some selections allow extra parameters for fine tuning the values generated. If extra parameters are allowed, an additional prompt will be displayed. See Faker Documentation for details on each type and extra parameters ('provider' in the Faker docs corresponds to 'type').

[Configure Page] Don't forget to click save and then move onto the next page. The user can always come back and update these parameters. [Train Page] On the next page, the user can fine-tune the training parameters for the selected model.

Each model has a different set of parameters (for example the neural-network-based models have an option for # of training epochs). Usually, the defaults are okay to start with. The user may want to come back and adjust if the results don't look right in the last step.

Once ready, click 'Train Model'. This may take only a few seconds (small data with a GaussianCopula model for example), or a few hours for some of the neural-net-based models on complex data. This will trigger a training task that runs in the background.

The refresh button will update the status displayed, or the next time the app is loaded, the status will be reflected on the [Start Page]. If the training succeeds, a notification will be displayed here. Click Next to continue.

[Generate Page] Select the number of rows to generate and click 'Generate Data'. This may take a few moments but is much faster than training.

Visit Website [email protected]
Intelligence Score
18/100
Your Rating
Platforms

1

Domain

datamynd.ai

Data Quality SignalsD

Freshness

Single-source

API Status

No API

44Completeness

Compliance (vendor-reported)

SOC 2
ISO 27001
GDPR
CCPA

Quality Breakdown

Data Coverage
20
Documentation
75
Compliance
0
Pricing Clarity
0

Signal Analysis

Categories

AI & ML

Listed On

snowflake_marketplace
Intelligence Scores3/7 pillars scored
Market Presence
12
Integration
20
Business Maturity
3
Extended Analytics

Similar Vendors

TA

Tonic.ai

1 product

80%
MA

Matillion

1 product

80%

About DataMynd

DataMynd is an alternative data vendor. DataMynd specializes in ai & ml data. This vendor has a Vedex Intelligence Score of 18 out of 100, reflecting market presence, compliance posture, integration readiness, and business maturity.

Data Categories

DataMynd operates in the following alternative data categories.

AI & ML

Related Resources

View Due Diligence ReportBrowse All VendorsBrowse All Products